Environment
- 3 combined LINSTOR nodes + diskless clients (Kubernetes nodes via Piraeus). LINSTOR 1.33.1, Ubuntu 22.04, kernel 6.8.0-106-generic.
- node1: DRBD 9.3.2, drbd-utils 9.34.0. node2: DRBD 9.3.0, drbd-utils 9.34.0-rc.1. Kubernetes diskless nodes: DRBD 9.2.14 (Piraeus build).
- node1-node2 node-connection:
DrbdOptions/Net/load-balance-paths yes (lb-tcp), 2 paths over 2 dedicated NICs (MTU 9000, no NIC errors).
1. Kernel Oops when changing the transport of a connected resource (lb-tcp -> tcp)
linstor resource-connection drbd-peer-options --load-balance-paths no node1 node2 <res> with both satellites online; resource Secondary on node1, Primary on node2, connected over lb-tcp. Both sides re-created the connection at the same time. node1 kernel (UTC):
11:23:24.208 remote state change from node2 -> conn( Connected -> TearDown ) -> Unconnected -> Connecting
11:23:24.314 conn( Connecting -> Disconnecting ) [del-peer]; "Failed to initiate connection, err=-512"
11:23:24.346 StandAlone -> Unconnected [connect] -> Connecting
11:23:24.346 BUG: unable to handle page fault -> Oops #1 in the receiver thread:
native_queued_spin_lock_slowpath <- _raw_spin_lock_irqsave <- prepare_to_wait_event
<- dtt_wait_for_connect [drbd_transport_tcp] <- dtt_connect <- drbd_receiver
"exited with irqs disabled", "exited with preempt_count 1"
11:23:24.851 "lb-tcp:node2: Error receiving initial packet. err = 65534"
11:23:26.099 Oops #2 (kworker): NULL pointer dereference in dtl_socket_ok_or_free <- dtl_accept_work_fn [drbd_transport_lb_tcp]
Changing the same option while node1 was down for the resource worked fine (5 resources).
2. After the Oops, down of that resource never completes
On reboot, drbd-graceful-shutdown (down all) blocked for more than 9 minutes on that resource. Every other resource, including a Primary one, went down cleanly first. During the hang the peer kept completing TCP handshakes to the resource port and logged timeout while waiting for feature packet every ~30 s (listener alive, receiver thread dead). Only a BMC reset recovered the node.
3. lb-tcp connection drops under load (also with 9.3.0 on both sides)
- The busiest resources (iSCSI-exported VMFS LUNs) drop the node1<->node2 connection 1-3 times per day since at least 2026-09-24.
- node1:
sock was shut down by peer while receiving data -> BrokenPipe, error receiving P_DATA, e: -5 l: 4096!.
- node2:
meta connection shut down by peer or lb-tcp:node1: _dtl_send_page: size=4096 len=4096 sent=-32 -> NetworkFailure.
- ~80 lighter resources on the same lb-tcp node-connection: 0 drops in the same period. The LUN resource group uses max-buffers 36864 and sndbuf/rcvbuf 1 MiB.
- After moving only the node1-node2 connection of the LUNs to tcp: 0 drops so far (overnight + day).
4. Inconsistent replica never becomes UpToDate when the Primary is a 9.2.14 diskless node
3 resources: diskful on node1 and node2, Primary is a diskless Kubernetes node running 9.2.14. The node1 replica was created on 2026-10-05 and has been Inconsistent since then.
- With the diskless Primary connected:
repl( Established -> WFBitMapT ) [diskless-primary] -> Began resync as SyncTarget (will sync 0 KB [0 bits set]) -> stays SyncTarget forever (out-of-sync:0, rs-in-flight:0 on both sides).
- Deleting and re-creating the node1 replica (fresh LV, new connection) reproduces it: 21384 KB [5346 bits] are synced, then the same stuck SyncTarget.
- Disconnecting the diskless Primary from node1 immediately completes the stuck resync (
Resync done (total 254 sec ...), repl -> Established), but node1's disk stays Inconsistent.
drbdadm invalidate on node1 while the Primary is disconnected: full resync from node2, 15729336 KB [3932334 bits] in 169 s, Resync done, repl Established, and the disk is STILL Inconsistent (no disk( Inconsistent -> UpToDate )).
- Reconnecting the Primary:
uuid_compare()=target-use-bitmap by rule=bitmap-peer -> [diskless-primary] -> 0-bit resync stuck again.
- In the same hour, ~77 other resources on the same node pair, whose node1 replica was Outdated rather than Inconsistent, resynced to UpToDate normally. The diskless Primary negotiates protocol 123 with feature flags 0x7f; node2 negotiates 0x1ff.
Possibly the same root cause as #148 (resync stays Established/Inconsistent after the sync source loses a diskless primary), seen here with 9.3.0/9.3.2 diskful nodes and a 9.2.14 diskless primary.
Questions
- Is the Oops in
dtt_wait_for_connect / dtl_accept_work_fn (del-peer racing connect) known, and which 9.3.x release fixes it?
- Is the lb-tcp drop under load (
sent=-32 / shut down by peer) known? Is plain tcp the recommendation for heavily loaded resources?
- Is an Inconsistent SyncTarget that never reaches UpToDate with a 9.2.x diskless Primary a known mixed-version issue? What is the supported way to bring that replica to UpToDate?
node1 kernel log around the Oops (sanitized)
11:23:24.208357 node1 kernel: drbd res1: Preparing remote state change 2324492637: 1->0 conn( Disconnecting )
11:23:24.208543 node1 kernel: drbd res1/0 drbd1001: Would lose quorum, but using tiebreaker logic to keep
11:23:24.208611 node1 kernel: drbd res1 node2: Committing remote state change 2324492637 (primary_nodes=0)
11:23:24.208912 node1 kernel: drbd res1 node2: conn( Connected -> TearDown ) peer( Secondary -> Unknown ) [remote]
11:23:24.208974 node1 kernel: drbd res1/0 drbd1001 node2: pdsk( UpToDate -> DUnknown ) repl( Established -> Off ) [remote]
11:23:24.209032 node1 kernel: drbd res1 node2: Terminating sender thread
11:23:24.209088 node1 kernel: drbd res1 node2: Starting sender thread (peer-node-id 1)
11:23:24.247388 node1 kernel: drbd res1 node2: Connection closed
11:23:24.247516 node1 kernel: drbd res1 node2: helper command: /sbin/drbdadm disconnected
11:23:24.260350 node1 kernel: drbd res1 node2: helper command: /sbin/drbdadm disconnected exit code 0
11:23:24.260465 node1 kernel: drbd res1 node2: conn( TearDown -> Unconnected ) [disconnected]
11:23:24.260529 node1 kernel: drbd res1 node2: Restarting receiver thread
11:23:24.260582 node1 kernel: drbd res1 node2: conn( Unconnected -> Connecting ) [connecting]
11:23:24.314308 node1 kernel: drbd res1 node2: conn( Connecting -> Disconnecting ) [del-peer]
11:23:24.314398 node1 kernel: drbd res1 node2: Failed to initiate connection, err=-512
11:23:24.314454 node1 kernel: drbd res1 node2: Terminating sender thread
11:23:24.314487 node1 kernel: drbd res1 node2: Starting sender thread (peer-node-id 1)
11:23:24.316353 node1 kernel: drbd res1 node2: Connection closed
11:23:24.316465 node1 kernel: drbd res1 node2: helper command: /sbin/drbdadm disconnected
11:23:24.329340 node1 kernel: drbd res1 node2: helper command: /sbin/drbdadm disconnected exit code 0
11:23:24.329422 node1 kernel: drbd res1 node2: conn( Disconnecting -> StandAlone ) [disconnected]
11:23:24.329466 node1 kernel: drbd res1 node2: Terminating receiver thread
11:23:24.329501 node1 kernel: drbd res1 node2: Terminating sender thread
11:23:24.340346 node1 kernel: drbd res1 node2: Starting sender thread (peer-node-id 1)
11:23:24.346340 node1 kernel: drbd res1 node2: conn( StandAlone -> Unconnected ) [connect]
11:23:24.346416 node1 kernel: drbd res1 node2: Starting receiver thread (peer-node-id 1)
11:23:24.346457 node1 kernel: drbd res1 node2: conn( Unconnected -> Connecting ) [connecting]
11:23:24.346492 node1 kernel: BUG: unable to handle page fault for address: ffffffff9aaaa7c0
11:23:24.481747 node1 kernel: Oops: 0002 [#1] PREEMPT SMP PTI
11:23:24.481829 node1 kernel: CPU: 16 PID: 120861 Comm: drbd_r_res1 Tainted: G OE 6.8.0-106-generic #106~22.04.1-Ubuntu
11:23:24.481870 node1 kernel: Hardware name: Supermicro PIO-648R-E1CR36L+-ST031/X10DRi-T4+, BIOS 3.3 10/24/2020
11:23:24.481904 node1 kernel: RIP: 0010:native_queued_spin_lock_slowpath+0x25f/0x300
11:23:24.482323 node1 kernel: Call Trace:
11:23:24.482793 node1 kernel: __raw_spin_lock_irqsave+0x57/0x80
11:23:24.483109 node1 kernel: _raw_spin_lock_irqsave+0xe/0x20
11:23:24.483284 node1 kernel: prepare_to_wait_event+0x1b/0x130
11:23:24.483569 node1 kernel: dtt_wait_for_connect.constprop.0+0x37e/0x5f0 [drbd_transport_tcp]
11:23:24.483938 node1 kernel: dtt_connect+0x3d6/0xf9e [drbd_transport_tcp]
11:23:24.484071 node1 kernel: drbd_receiver+0xda/0xae0 [drbd]
11:23:24.484411 node1 kernel: drbd_thread_setup+0xeb/0x2d0 [drbd]
11:23:24.484720 node1 kernel: kthread+0xf2/0x120
11:23:24.484988 node1 kernel: ret_from_fork+0x47/0x70
11:23:24.485259 node1 kernel: ret_from_fork_asm+0x1b/0x30
11:23:24.485630 node1 kernel: Modules linked in: bcache dm_cache dm_writecache nvme_rdma nvme_fabrics nvmet_rdma nvmet nvme_keyring nvme_core nvme_auth vhost_vsock vmw_vsock_virtio_transport_common vhost vhost_iotlb vsock drbd_tra
11:23:24.485975 node1 kernel: raid0 mlx4_ib ib_uverbs ib_core mlx4_en hid_generic usbhid hid dm_thin_pool dm_persistent_data dm_bio_prison dm_bufio libcrc32c ses enclosure scsi_transport_sas i2c_i801 crct10dif_pclmul crc32_pclmul
11:23:24.486630 node1 kernel: RIP: 0010:native_queued_spin_lock_slowpath+0x25f/0x300
11:23:24.488112 node1 kernel: note: drbd_r_res1[120861] exited with irqs disabled
11:23:24.488235 node1 kernel: note: drbd_r_res1[120861] exited with preempt_count 1
11:23:24.851360 node1 kernel: drbd res1 lb-tcp:node2: Error receiving initial packet. err = 65534
11:23:26.099389 node1 kernel: BUG: kernel NULL pointer dereference, address: 0000000000000012
11:23:26.106199 node1 kernel: Oops: 0000 [#2] PREEMPT SMP PTI
11:23:26.106313 node1 kernel: CPU: 9 PID: 116101 Comm: kworker/9:2 Tainted: G D OE 6.8.0-106-generic #106~22.04.1-Ubuntu
11:23:26.106387 node1 kernel: Hardware name: Supermicro PIO-648R-E1CR36L+-ST031/X10DRi-T4+, BIOS 3.3 10/24/2020
11:23:26.106516 node1 kernel: RIP: 0010:dtl_socket_ok_or_free.isra.0+0x20/0x90 [drbd_transport_lb_tcp]
11:23:26.107125 node1 kernel: Call Trace:
11:23:26.107216 node1 kernel: dtl_accept_work_fn+0x306/0x730 [drbd_transport_lb_tcp]
11:23:26.107271 node1 kernel: process_one_work+0x184/0x3a0
11:23:26.107338 node1 kernel: worker_thread+0x306/0x440
11:23:26.107507 node1 kernel: kthread+0xf2/0x120
11:23:26.107608 node1 kernel: ret_from_fork+0x47/0x70
11:23:26.107699 node1 kernel: ret_from_fork_asm+0x1b/0x30
11:23:26.107791 node1 kernel: Modules linked in: bcache dm_cache dm_writecache nvme_rdma nvme_fabrics nvmet_rdma nvmet nvme_keyring nvme_core nvme_auth vhost_vsock vmw_vsock_virtio_transport_common vhost vhost_iotlb vsock drbd_tra
11:23:26.107920 node1 kernel: raid0 mlx4_ib ib_uverbs ib_core mlx4_en hid_generic usbhid hid dm_thin_pool dm_persistent_data dm_bio_prison dm_bufio libcrc32c ses enclosure scsi_transport_sas i2c_i801 crct10dif_pclmul crc32_pclmul
11:23:26.108080 node1 kernel: RIP: 0010:native_queued_spin_lock_slowpath+0x25f/0x300
11:23:26.108640 node1 kernel: note: kworker/9:2[116101] exited with irqs disabled
Environment
DrbdOptions/Net/load-balance-paths yes(lb-tcp), 2 paths over 2 dedicated NICs (MTU 9000, no NIC errors).1. Kernel Oops when changing the transport of a connected resource (lb-tcp -> tcp)
linstor resource-connection drbd-peer-options --load-balance-paths no node1 node2 <res>with both satellites online; resource Secondary on node1, Primary on node2, connected over lb-tcp. Both sides re-created the connection at the same time. node1 kernel (UTC):Changing the same option while node1 was down for the resource worked fine (5 resources).
2. After the Oops,
downof that resource never completesOn reboot,
drbd-graceful-shutdown(down all) blocked for more than 9 minutes on that resource. Every other resource, including a Primary one, went down cleanly first. During the hang the peer kept completing TCP handshakes to the resource port and loggedtimeout while waiting for feature packetevery ~30 s (listener alive, receiver thread dead). Only a BMC reset recovered the node.3. lb-tcp connection drops under load (also with 9.3.0 on both sides)
sock was shut down by peer while receiving data-> BrokenPipe,error receiving P_DATA, e: -5 l: 4096!.meta connection shut down by peerorlb-tcp:node1: _dtl_send_page: size=4096 len=4096 sent=-32-> NetworkFailure.4. Inconsistent replica never becomes UpToDate when the Primary is a 9.2.14 diskless node
3 resources: diskful on node1 and node2, Primary is a diskless Kubernetes node running 9.2.14. The node1 replica was created on 2026-10-05 and has been Inconsistent since then.
repl( Established -> WFBitMapT ) [diskless-primary]->Began resync as SyncTarget (will sync 0 KB [0 bits set])-> stays SyncTarget forever (out-of-sync:0, rs-in-flight:0 on both sides).Resync done (total 254 sec ...), repl -> Established), but node1's disk stays Inconsistent.drbdadm invalidateon node1 while the Primary is disconnected: full resync from node2, 15729336 KB [3932334 bits] in 169 s,Resync done, repl Established, and the disk is STILL Inconsistent (nodisk( Inconsistent -> UpToDate )).uuid_compare()=target-use-bitmap by rule=bitmap-peer->[diskless-primary]-> 0-bit resync stuck again.Possibly the same root cause as #148 (resync stays Established/Inconsistent after the sync source loses a diskless primary), seen here with 9.3.0/9.3.2 diskful nodes and a 9.2.14 diskless primary.
Questions
dtt_wait_for_connect/dtl_accept_work_fn(del-peer racing connect) known, and which 9.3.x release fixes it?sent=-32/shut down by peer) known? Is plain tcp the recommendation for heavily loaded resources?node1 kernel log around the Oops (sanitized)