== 1. Summary ==
On a 4-node LINSTOR/DRBD cluster with the RDMA transport, listener cm_ids are
leaked during connection establishment. A leaked cm_id stays in state LISTEN on
its TCP port for the lifetime of the boot. Because LINSTOR re-uses the lowest
free port number from its auto-range, the port is later handed to a new
resource, and every connect attempt of that resource fails permanently with:
drbd rdma:: rdma_bind_addr error -98
drbd : Failed to initiate connection, err=-98
drbd : conn( Connecting -> Disconnecting )
drbd : conn( Disconnecting -> StandAlone ) [disconnected]
Affected resources go StandAlone on both sides and stay unreplicated. Only a
reboot recovers the port. We first saw this on 9.3.2, upgraded to 9.3.3 on
2026-09-25, and the behaviour is unchanged.
We believe this is the exact defect you fixed in the new rdma2 transport with
commit 0f9f1e3 ("drbd: rdma2: release the passive connect state when
establishing fails", 2026-09-25), whose own commit message ends with:
"Inherited from drbd_transport_rdma.c, which has the same branch."
We are asking for that fix to be backported to drbd_transport_rdma.c in 9.3.x.
Details and our verification are in section 6.
== 2. Environment ==
DRBD module 9.3.3 (api:2/proto:118-124), DKMS build 2026-09-25 09:33
GIT-hash: 97da760
Transport rdma (9.3.3) -- "Transports (api:22)" from /proc/drbd
LINSTOR controller 1.33.3, satellites in podman containers
OS Ubuntu 22.04.5 LTS
Kernel 6.8.0-90-generic (x86_64)
RDMA stack MLNX_OFED_LINUX-24.10-4.1.4.0
HCA Mellanox MT4125 (ConnectX-6 Lx), fw 22.43.4100, mlx5_ib/mlx5_core
Fabric RoCE, mlx5_0/1 ACTIVE / LINK_UP (mlx5_1 unused, DOWN)
Nodes 4 storage nodes, 2 LINSTOR controllers, diskless clients
Scale 87 DRBD resources, 176 connections on the node below
Workload Kubernetes PVCs: dozens of resource creations and deletions
per day; the failure correlates with delete-under-connect
The port range LINSTOR allocates from is TcpPortAutoRange = 11000-13000, and it
always picks the lowest free number.
== 3. Impact ==
Per node, right now (uptime 3 days since the 9.3.3 reboot):
node-A (192.0.2.34) 7 poisoned ports, 7 resources StandAlone (14 connections)
node-B same 7, peer side, all Connecting
node-C 4 poisoned ports, 3 resources StandAlone
node-D same 4, peer side, all Connecting
(node-A and node-B are one replica pair in one datacentre, node-C and node-D
the other; node-A and node-C also run the LINSTOR controllers. All four are
identical hardware and configuration.)
The affected volumes are mounted and serving I/O, but they run without a
replica. The number of unusable ports grows monotonically until reboot.
== 4. Evidence ==
--- 4.1 The leaked listeners are there, and they own no resource ---
rdma res show cm_id | grep LISTEN
link mlx5_0/1 cm-idn 1914 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11057 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 5840 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11059 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 10028 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11058 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 10232 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11060 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 13153 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11023 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 18486 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11056 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 35524 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11055 dst-addr 0.0.0.0:0
Each of these 7 port numbers is now assigned by LINSTOR to a resource that is
StandAlone and cannot connect:
11023 -> pvc-c72bce27-94a3-4f91-85db-a11f7f8c4e02
11055 -> pvc-e613451b-e92f-483f-a53c-e58b30bd1bce
11056 -> pvc-542523db-9026-4fa5-bf60-acbb97bf3b31
11057 -> pvc-ce1bfe0b-9cf2-4c91-baf3-f92d421cc71a
11058 -> pvc-678c191e-44b4-4cc1-94e2-d4b514e89b81
11059 -> pvc-1b7d50ba-c090-4cda-85fb-3fb2b5ca04ef
11060 -> pvc-96a9045d-d118-421b-bc7d-26cb86c06bb4
The listening cm_id does not belong to these resources: their own bind is what
returns -98, so they never obtained a listener. The listeners are left over
from earlier, already deleted resources that used the same port numbers.
We verified the leaked cm_id outlives the complete removal of its resource: on
2026-09-14 (on 9.3.2) ports 11005 and 11007 were still in LISTEN after the
corresponding resources had been fully deleted -- zvol gone, .res file gone, no
resource definition left in LINSTOR -- yet the RDMA-CM sockets were alive. So
nothing in DRBD or LINSTOR owns those cm_ids any more.
drbdsetup status --json | (connection-state histogram)
Connected 162, StandAlone 14 <- 7 resources x 2 connections
--- 4.2 The error is continuous, not a one-off ---
journalctl -k --since -6h | grep -c 'rdma_bind_addr error -98'
1135
2026-09-28T09:10:22 kernel: drbd pvc-c72bce27-... rdma:client-1: rdma_bind_addr error -98
2026-09-28T09:10:22 kernel: drbd pvc-c72bce27-... client-1: Failed to initiate connection, err=-98
2026-09-28T09:10:22 kernel: drbd pvc-c72bce27-... rdma:node-B: rdma_bind_addr error -98
2026-09-28T09:10:22 kernel: drbd pvc-c72bce27-... node-B: Failed to initiate connection, err=-98
2026-09-28T09:10:38 kernel: (same four lines, next retry)
Note the bind fails for every peer of the resource at once (replica peer and
diskless client alike): it is the local bind that is refused, not a per-peer
address conflict.
--- 4.3 Accumulation from a clean boot on 9.3.3 ---
All four nodes were rebooted on 2026-09-25 09:25-09:57 with the freshly built
9.3.3 module, so the count starts from zero. Poisoned ports on node-A, from
periodic snapshots:
+12 h 11057, 11059
+18 h 11023, 11057, 11058, 11059, 11060
+24 h + 11056
+30 h + 11055 -> 7 ports
+36..66 h the same 7 (plateau)
On node-C the first port (11001) was already poisoned 2 hours after boot;
4 ports by +32 h.
The plateau is not the leak stopping. LINSTOR hands out the lowest free port,
the poisoned numbers are the lowest free ones, so every new resource gets a
poisoned port, fails with -98, is deleted, and frees the number for the next
one. Within 12 hours of retained kernel log we counted 2426 "rdma_bind_addr
error -98" messages from 247 distinct resource names on those 7 ports.
--- 4.4 The rate is the same as on 9.3.2 ---
Counted from kernel logs on node-C:
9.3.2 2026-09-17 13:07 -> 23:21 (10.2 h) 4397 msgs / 276 resources ~430/h
9.3.2 2026-09-17 19:02 -> 18 05:22 (10.3 h) 4614 msgs / 293 resources ~448/h
9.3.3 2026-09-27 20:37 -> 28 06:10 ( 9.6 h) 3224 msgs / 198 resources ~338/h
The difference is within the variation of our PVC delete load. 9.3.3 did not
change this behaviour.
== 5. What we ruled out ==
- Fabric / configuration: on the same node, over the same addresses and the
same transport, other resources are Connected with peer-disk UpToDate. The
.res files match byte for byte on both sides (node-id, port, shared-secret,
transport "rdma"). PrefNic is set on all storage nodes.
- Port double-allocation by LINSTOR: the ports LINSTOR assigns are unique
across resource definitions; the conflict is with a leaked kernel cm_id,
not with another live resource.
- Split-brain / data divergence: excluded, the resources never connected.
- Reconnect as a remedy: "drbdadm connect" returns rc=0 and then immediately
goes bind -> EADDRINUSE -> Disconnecting -> StandAlone. Restarting the
LINSTOR satellite changes nothing; the state is in the kernel.
- Process/PID limits, MTU, RoCE v1 vs v2: all checked and not involved.
Related earlier incident (2026-09-08, on 9.3.2, same cluster): a
"drbdsetup del-peer 2" was stuck in D state for 6 days with
[<0>] _drbd_thread_stop+0xf7/0x120 [drbd]
[<0>] del_connection+0x84/0x1b0 [drbd]
[<0>] adm_disconnect.constprop.0+0x1be/0x290 [drbd]
[<0>] drbd_adm_del_peer+0x16/0x20 [drbd]
The resource had been created 7 minutes earlier and was being deleted while the
connection was still being established. That hang has not recurred on 9.3.3 in
3 days (no task in D, no "blocked for more than", no ASSERT FAILED), which is
consistent with b0a0fd9 being in 9.3.3. The listener leak, however, is
still fully reproducible.
== 6. Analysis: dtr_path_established_work_fn() leaves PCS_FINISHING behind ==
We read drbd_transport_rdma.c from the drbd-9.3.3 and drbd-9.3.4 tags. In
dtr_path_established_work_fn() the function takes the passive connect state
from PCS_CONNECTING to PCS_FINISHING and is supposed to release it, together
with the listener, at the "out:" label. When the path is no longer OK after the
first flow-control message, it jumps to "out_put:" and skips that release
(drbd-9.3.3 lines 857-915, drbd-9.3.4 lines 812-870):
p = atomic_cmpxchg(&cs->passive_state, PCS_CONNECTING, PCS_FINISHING);
if (p < PCS_CONNECTING)
goto out;
...
err = dtr_send_flow_control_msg(path, GFP_NOIO);
...
schedule_timeout(HZ / 4);
if (!dtr_path_ok(path)) {
if (path->cs.active)
dtr_cma_retry_connect(path, path->cm);
goto out_put; /* <-- skips "out:" */
}
...
out:
atomic_set(&cs->active_state, PCS_INACTIVE);
p = atomic_xchg(&cs->passive_state, PCS_INACTIVE);
if (p > PCS_INACTIVE)
drbd_put_listener(&path->path); /* the port is freed here /
wake_up(&cs->wq);
out_put:
kref_put(&cm->kref, dtr_destroy_cm); / for work */
Consequences, both of which we observe:
- drbd_put_listener() is never called, so the listener keeps its cm_id in
LISTEN on that port until reboot -> rdma_bind_addr -98 for whoever gets
that port number next;
- passive_state stays PCS_FINISHING, so every later __dtr_disconnect_path()
waits 60 s for PCS_INACTIVE (the receiver among them, via
dtr_finish_connect()), which matches the slow, timing-out disconnects we
see around the failing resources.
This is precisely what your commit 0f9f1e3 fixes in rdma2, and its commit
message says the same branch exists in drbd_transport_rdma.c. We also confirmed
this branch is unchanged in drbd-9.3.4: between the drbd-9.3.3 and drbd-9.3.4
tags only three commits touch drbd_transport_rdma.c, and all three are
cosmetic (73716fe "restructure to eliminate forward declarations",
57109d5 "replace bitfield struct with u32", e04342e "assorted style
cleanups"). So upgrading to 9.3.4 would not help us.
For completeness, the cm-kref fixes that are already in our 9.3.3 and are
therefore not the answer here: 6e726de, 4506ab1, b0a0fd9,
2d2343b.
== 7. What we are asking for ==
(a) Please backport 0f9f1e3 to drbd_transport_rdma.c and release it in a
9.3.x build. The rdma2 patch applies to the same branch in our transport:
if (!dtr_path_ok(path)) {
if (path->cs.active)
dtr_cma_retry_connect(path, path->cm);
+ /* Release the PCS_FINISHING hold taken above, as the success path
+ * does below, or __dtr_disconnect_path() waits a minute for it.
+ */
+ p = atomic_xchg(&cs->passive_state, PCS_INACTIVE);
+ if (p > PCS_INACTIVE)
+ drbd_put_listener(&path->path);
+ wake_up(&cs->wq);
goto out_put;
}
(b) If you prefer, we can run a test build for you. This cluster reproduces the
leak within 2-12 hours of a reboot under normal production load, so it is a
fast and realistic verification environment. We are willing to run a
pre-release or debug-instrumented module on one node pair.
== 8. Questions ==
- Do you agree that the branch above is the (or a) source of our leaked
LISTEN cm_ids, and is a backport to 9.3.x planned?
- Beyond 0f9f1e3, does drbd_transport_rdma.c need the rest of the rdma2
hardening -- 35f7a33 ("never destroy a cm_id from inside its own
event callback"), 3f95752 ("retry the connect when a signal
interrupts it"), 1f91341 ("add a connect-attempt watchdog")? We would
rather not cherry-pick these ourselves without your assessment.
- What is the intended status of drbd_transport_rdma2 -- is it going to
replace drbd_transport_rdma in a 9.3.x or 9.4 release, and would you
recommend we evaluate it for this workload?
- Is there any supported way to release a leaked listener cm_id without
rebooting the node (a debugfs knob, a drbdsetup command, an unload/reload
sequence that is safe with live resources)? Today a reboot is our only
recovery, and it costs a full resync of everything on the node.
- Do you see the same class of bug in the code path we cannot read from
here -- for reference, 82e8fa8 ("lb-tcp: drop the listeners of the
paths that lost the connect race", 2026-09-17) looks like the same defect
class fixed only in lb-tcp.
== 9. Mitigations available to us ==
- Rolling reboots to reclaim the ports. This is the only thing we have
actually done so far (all four nodes on 2026-09-25, together with the
upgrade to 9.3.3); the ports started leaking again within hours.
- Moving an individual affected replica pair to transport "tcp" should bring
it up immediately: the poisoned port numbers are free in the TCP socket
namespace ("ss -tanl" shows no listener on them -- only RDMA-CM holds them),
and drbd_transport_tcp.ko is built and available. Not tried yet, since it
silently drops RoCE for that pair.
- Shifting LINSTOR's TcpPortAutoRange above the poisoned numbers would stop
new resources from inheriting them, at the cost of leaving the leaked ports
behind until the next reboot. Not applied yet.
None of these addresses the leak itself, which is why we are opening this case.
== 1. Summary ==
On a 4-node LINSTOR/DRBD cluster with the RDMA transport, listener cm_ids are
leaked during connection establishment. A leaked cm_id stays in state LISTEN on
its TCP port for the lifetime of the boot. Because LINSTOR re-uses the lowest
free port number from its auto-range, the port is later handed to a new
resource, and every connect attempt of that resource fails permanently with:
drbd rdma:: rdma_bind_addr error -98
drbd : Failed to initiate connection, err=-98
drbd : conn( Connecting -> Disconnecting )
drbd : conn( Disconnecting -> StandAlone ) [disconnected]
Affected resources go StandAlone on both sides and stay unreplicated. Only a
reboot recovers the port. We first saw this on 9.3.2, upgraded to 9.3.3 on
2026-09-25, and the behaviour is unchanged.
We believe this is the exact defect you fixed in the new rdma2 transport with
commit 0f9f1e3 ("drbd: rdma2: release the passive connect state when
establishing fails", 2026-09-25), whose own commit message ends with:
"Inherited from drbd_transport_rdma.c, which has the same branch."
We are asking for that fix to be backported to drbd_transport_rdma.c in 9.3.x.
Details and our verification are in section 6.
== 2. Environment ==
DRBD module 9.3.3 (api:2/proto:118-124), DKMS build 2026-09-25 09:33
GIT-hash: 97da760
Transport rdma (9.3.3) -- "Transports (api:22)" from /proc/drbd
LINSTOR controller 1.33.3, satellites in podman containers
OS Ubuntu 22.04.5 LTS
Kernel 6.8.0-90-generic (x86_64)
RDMA stack MLNX_OFED_LINUX-24.10-4.1.4.0
HCA Mellanox MT4125 (ConnectX-6 Lx), fw 22.43.4100, mlx5_ib/mlx5_core
Fabric RoCE, mlx5_0/1 ACTIVE / LINK_UP (mlx5_1 unused, DOWN)
Nodes 4 storage nodes, 2 LINSTOR controllers, diskless clients
Scale 87 DRBD resources, 176 connections on the node below
Workload Kubernetes PVCs: dozens of resource creations and deletions
per day; the failure correlates with delete-under-connect
The port range LINSTOR allocates from is TcpPortAutoRange = 11000-13000, and it
always picks the lowest free number.
== 3. Impact ==
Per node, right now (uptime 3 days since the 9.3.3 reboot):
node-A (192.0.2.34) 7 poisoned ports, 7 resources StandAlone (14 connections)
node-B same 7, peer side, all Connecting
node-C 4 poisoned ports, 3 resources StandAlone
node-D same 4, peer side, all Connecting
(node-A and node-B are one replica pair in one datacentre, node-C and node-D
the other; node-A and node-C also run the LINSTOR controllers. All four are
identical hardware and configuration.)
The affected volumes are mounted and serving I/O, but they run without a
replica. The number of unusable ports grows monotonically until reboot.
== 4. Evidence ==
--- 4.1 The leaked listeners are there, and they own no resource ---
rdma res show cm_id | grep LISTEN
link mlx5_0/1 cm-idn 1914 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11057 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 5840 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11059 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 10028 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11058 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 10232 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11060 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 13153 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11023 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 18486 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11056 dst-addr 0.0.0.0:0
link mlx5_0/1 cm-idn 35524 state LISTEN ps TCP comm [drbd_transport_r src-addr 192.0.2.34:11055 dst-addr 0.0.0.0:0
Each of these 7 port numbers is now assigned by LINSTOR to a resource that is
StandAlone and cannot connect:
11023 -> pvc-c72bce27-94a3-4f91-85db-a11f7f8c4e02
11055 -> pvc-e613451b-e92f-483f-a53c-e58b30bd1bce
11056 -> pvc-542523db-9026-4fa5-bf60-acbb97bf3b31
11057 -> pvc-ce1bfe0b-9cf2-4c91-baf3-f92d421cc71a
11058 -> pvc-678c191e-44b4-4cc1-94e2-d4b514e89b81
11059 -> pvc-1b7d50ba-c090-4cda-85fb-3fb2b5ca04ef
11060 -> pvc-96a9045d-d118-421b-bc7d-26cb86c06bb4
The listening cm_id does not belong to these resources: their own bind is what
returns -98, so they never obtained a listener. The listeners are left over
from earlier, already deleted resources that used the same port numbers.
We verified the leaked cm_id outlives the complete removal of its resource: on
2026-09-14 (on 9.3.2) ports 11005 and 11007 were still in LISTEN after the
corresponding resources had been fully deleted -- zvol gone, .res file gone, no
resource definition left in LINSTOR -- yet the RDMA-CM sockets were alive. So
nothing in DRBD or LINSTOR owns those cm_ids any more.
drbdsetup status --json | (connection-state histogram)
Connected 162, StandAlone 14 <- 7 resources x 2 connections
--- 4.2 The error is continuous, not a one-off ---
journalctl -k --since -6h | grep -c 'rdma_bind_addr error -98'
1135
2026-09-28T09:10:22 kernel: drbd pvc-c72bce27-... rdma:client-1: rdma_bind_addr error -98
2026-09-28T09:10:22 kernel: drbd pvc-c72bce27-... client-1: Failed to initiate connection, err=-98
2026-09-28T09:10:22 kernel: drbd pvc-c72bce27-... rdma:node-B: rdma_bind_addr error -98
2026-09-28T09:10:22 kernel: drbd pvc-c72bce27-... node-B: Failed to initiate connection, err=-98
2026-09-28T09:10:38 kernel: (same four lines, next retry)
Note the bind fails for every peer of the resource at once (replica peer and
diskless client alike): it is the local bind that is refused, not a per-peer
address conflict.
--- 4.3 Accumulation from a clean boot on 9.3.3 ---
All four nodes were rebooted on 2026-09-25 09:25-09:57 with the freshly built
9.3.3 module, so the count starts from zero. Poisoned ports on node-A, from
periodic snapshots:
+12 h 11057, 11059
+18 h 11023, 11057, 11058, 11059, 11060
+24 h + 11056
+30 h + 11055 -> 7 ports
+36..66 h the same 7 (plateau)
On node-C the first port (11001) was already poisoned 2 hours after boot;
4 ports by +32 h.
The plateau is not the leak stopping. LINSTOR hands out the lowest free port,
the poisoned numbers are the lowest free ones, so every new resource gets a
poisoned port, fails with -98, is deleted, and frees the number for the next
one. Within 12 hours of retained kernel log we counted 2426 "rdma_bind_addr
error -98" messages from 247 distinct resource names on those 7 ports.
--- 4.4 The rate is the same as on 9.3.2 ---
Counted from kernel logs on node-C:
9.3.2 2026-09-17 13:07 -> 23:21 (10.2 h) 4397 msgs / 276 resources ~430/h
9.3.2 2026-09-17 19:02 -> 18 05:22 (10.3 h) 4614 msgs / 293 resources ~448/h
9.3.3 2026-09-27 20:37 -> 28 06:10 ( 9.6 h) 3224 msgs / 198 resources ~338/h
The difference is within the variation of our PVC delete load. 9.3.3 did not
change this behaviour.
== 5. What we ruled out ==
same transport, other resources are Connected with peer-disk UpToDate. The
.res files match byte for byte on both sides (node-id, port, shared-secret,
transport "rdma"). PrefNic is set on all storage nodes.
across resource definitions; the conflict is with a leaked kernel cm_id,
not with another live resource.
goes bind -> EADDRINUSE -> Disconnecting -> StandAlone. Restarting the
LINSTOR satellite changes nothing; the state is in the kernel.
Related earlier incident (2026-09-08, on 9.3.2, same cluster): a
"drbdsetup del-peer 2" was stuck in D state for 6 days with
[<0>] _drbd_thread_stop+0xf7/0x120 [drbd]
[<0>] del_connection+0x84/0x1b0 [drbd]
[<0>] adm_disconnect.constprop.0+0x1be/0x290 [drbd]
[<0>] drbd_adm_del_peer+0x16/0x20 [drbd]
The resource had been created 7 minutes earlier and was being deleted while the
connection was still being established. That hang has not recurred on 9.3.3 in
3 days (no task in D, no "blocked for more than", no ASSERT FAILED), which is
consistent with b0a0fd9 being in 9.3.3. The listener leak, however, is
still fully reproducible.
== 6. Analysis: dtr_path_established_work_fn() leaves PCS_FINISHING behind ==
We read drbd_transport_rdma.c from the drbd-9.3.3 and drbd-9.3.4 tags. In
dtr_path_established_work_fn() the function takes the passive connect state
from PCS_CONNECTING to PCS_FINISHING and is supposed to release it, together
with the listener, at the "out:" label. When the path is no longer OK after the
first flow-control message, it jumps to "out_put:" and skips that release
(drbd-9.3.3 lines 857-915, drbd-9.3.4 lines 812-870):
out:
atomic_set(&cs->active_state, PCS_INACTIVE);
p = atomic_xchg(&cs->passive_state, PCS_INACTIVE);
if (p > PCS_INACTIVE)
drbd_put_listener(&path->path); /* the port is freed here /
wake_up(&cs->wq);
out_put:
kref_put(&cm->kref, dtr_destroy_cm); / for work */
Consequences, both of which we observe:
LISTEN on that port until reboot -> rdma_bind_addr -98 for whoever gets
that port number next;
waits 60 s for PCS_INACTIVE (the receiver among them, via
dtr_finish_connect()), which matches the slow, timing-out disconnects we
see around the failing resources.
This is precisely what your commit 0f9f1e3 fixes in rdma2, and its commit
message says the same branch exists in drbd_transport_rdma.c. We also confirmed
this branch is unchanged in drbd-9.3.4: between the drbd-9.3.3 and drbd-9.3.4
tags only three commits touch drbd_transport_rdma.c, and all three are
cosmetic (73716fe "restructure to eliminate forward declarations",
57109d5 "replace bitfield struct with u32", e04342e "assorted style
cleanups"). So upgrading to 9.3.4 would not help us.
For completeness, the cm-kref fixes that are already in our 9.3.3 and are
therefore not the answer here: 6e726de, 4506ab1, b0a0fd9,
2d2343b.
== 7. What we are asking for ==
(a) Please backport 0f9f1e3 to drbd_transport_rdma.c and release it in a
9.3.x build. The rdma2 patch applies to the same branch in our transport:
(b) If you prefer, we can run a test build for you. This cluster reproduces the
leak within 2-12 hours of a reboot under normal production load, so it is a
fast and realistic verification environment. We are willing to run a
pre-release or debug-instrumented module on one node pair.
== 8. Questions ==
LISTEN cm_ids, and is a backport to 9.3.x planned?
hardening -- 35f7a33 ("never destroy a cm_id from inside its own
event callback"), 3f95752 ("retry the connect when a signal
interrupts it"), 1f91341 ("add a connect-attempt watchdog")? We would
rather not cherry-pick these ourselves without your assessment.
replace drbd_transport_rdma in a 9.3.x or 9.4 release, and would you
recommend we evaluate it for this workload?
rebooting the node (a debugfs knob, a drbdsetup command, an unload/reload
sequence that is safe with live resources)? Today a reboot is our only
recovery, and it costs a full resync of everything on the node.
here -- for reference, 82e8fa8 ("lb-tcp: drop the listeners of the
paths that lost the connect race", 2026-09-17) looks like the same defect
class fixed only in lb-tcp.
== 9. Mitigations available to us ==
actually done so far (all four nodes on 2026-09-25, together with the
upgrade to 9.3.3); the ports started leaking again within hours.
it up immediately: the poisoned port numbers are free in the TCP socket
namespace ("ss -tanl" shows no listener on them -- only RDMA-CM holds them),
and drbd_transport_tcp.ko is built and available. Not tried yet, since it
silently drops RoCE for that pair.
new resources from inheriting them, at the cost of leaving the leaked ports
behind until the next reboot. Not applied yet.
None of these addresses the leak itself, which is why we are opening this case.