Repository navigation
Rolling-upgrade structure skew makes a local __dbis__ seq cursor row undecodable (root cause behind harper-pro#352 second call site) #1307
Description
Activity
Root cause confirmed (live CDP forensics on 8tj)
It is not a shared-structure skew — the
__dbis__structure table is intact (id 11 ={seqId, nodes}). It's a metadata-prefix misparse driven byisRocksDB=falseon the__dbis__RecordEncoder.The failing
[seq, 1]value is 27 bytes (end=27):01 01 01 01 94 62 00 00 0e 00 00 40 00 00 00 01 | 4b cb<float64> 90 └──────────────── 16-byte metadata prefix ─────┘ └ struct-11 {seqId, nodes:[]} ┘The 11 bytes from offset 16 are a perfectly valid struct-11 record — decoding
buffer.subarray(16)yields{seqId: 1781538711832, nodes: []}. The record was written with a RocksDB metadata prefix: 8-byte local-timestamp, thenbuffer[8]=0x0e=ACTION_32_BIT→ 32-bit flags0e000040(hasHAS_NODE_ID) →nodeId=1→ real record at offset 16.But the
__dbis__decoder hasisRocksDB=false(per #1260's note,handleLocalTimeForGetsis deliberately never called onattributesDbi). Sodecode()takes the LMDB heuristic branch:metadataFlags = buffer[0] | (buffer[1]<<5)=0x01 | (0x01<<5)=0x21→ reads onlyHAS_RESIDENCY_ID(4 bytes) → lands at offset 6, not 16.super.decode(subarray(6,27))reads0x00(fixint 0) and leaves 20 bytes → "Data read, but end of buffer not reached" →null.Proof: decoding the exact 27-byte record on the live decoder returns
nullas-is, and{localTime, version, nodeId:1, size:27, value:{seqId,nodes}}whenisRocksDBis temporarily set totrue.So: write/read asymmetry
seqrows are written through the RocksDB metadata path (8-byte ts + 32-bit flags incl. nodeId), but read back through a decoder whoseisRocksDB=falsecan't strip that prefix. The fix must reconcile the two without breaking the non-RocksDB structure-load path (superGetStructures) that #1260 relies on forattributesDbi.Open sub-question: why only some nodes/records hit this (the heavy audit/timestamp prefix appears on the
seqwrite but not, e.g., the table-attribute catalog rows in the same store) — i.e., what makes aseqrow get the versioned RocksDB prefix. Tracking toward a fix.🤖 Claude Opus 4.8 (1M context), on Kris's behalf.
C verification: it's a module-global encode-state leak, not an isRocksDB read asymmetry
The 16-byte prefix on the
seqrecord is spurious — it should not be there at all:- The
__dbis__store is openeduseVersions=false(OpenDBIObject:useVersions = isPrimary). It should never emit a version/timestamp prefix. The table-attribute catalog rows in the same store confirm this — their first byte is the struct-ref0x40+id(≥64), no prefix, so they never trip thenextByte<32decode heuristic and decode fine. recordUpdateris wired only toprimaryStore(Table.ts:200), never todbisDb. Theseqwrite (Table.ts:490 dbisDb.put) is a raw store put that does not set the*NextEncodingmodule globals (RecordEncoder.ts:104-109) — it encodes with whatever those globals currently hold.recordUpdatersetstimestampNextEncoding(658) +metadataInNextEncoding(675) +nodeId(700) before the encode that consumes+resets them. For a delete (record === undefined), the consumingstore.putSyncat line 742-744 is skipped, so the globals stay set. The next rawdbisDb.put(the seq cursor) then encodes with the leaked metadata → spurious 8-byte-ts + 32-bit-flags(+nodeId) prefix.
The leaked
nodeId=1(= the leader) and timestamp in the captured prefix match a just-applied replicated record, not the cursor — consistent with the leak. Timing/workload-dependent (needs a delete-then-seq-write interleave during apply), which matches the observed intermittency.Fix direction (revised: not A)
A(make the__dbis__decoderisRocksDB-aware) would read the spurious prefix but store boguslocalTime/nodeId, and flips a flag #1260's structure-load path depends on. The root issue is the leak. Preferred fix:- Gate metadata-prefix emission on the encoder's actual versioning config so a
useVersions=falsestore (__dbis__) never emits a prefix regardless of leaked globals (robust against all leak windows; makesseqrows prefix-free like the attribute rows → decode cleanly withisRocksDB=false). - Optionally also reset the
*NextEncodingglobals onrecordUpdater's skip path (defense-in-depth for the delete window).
harper-pro#394 (skip the undecodable row) remains the resilience layer; this stops the row from being mis-encoded in the first place.
🤖 Claude Opus 4.8 (1M context), on Kris's behalf.
- The
Fix PR: #1308 (draft). Implements option 3 from the analysis — encode hook ignores+preserves staged metadata for non-versioned stores, unconditional global reset in
recordUpdater's finally, and explicituseVersions=falsemarking of the__dbis__encoder (lmdb/rocksdb don't forward the option). Cross-model reviewed (Codex+Gemini, two rounds). — Claude (Opus 4.8)- added a commit that references this issue
on Jun 16, 2026 Closed by #1308 (merged 2026-06-16) — the encoding bug confirmed on a live node and fixed. The leaking module-global state that corrupted cursor rows is patched.
- added a commit that references this issue
on Jul 31, 2026 - added a commit that references this issue
on Aug 18, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
During an in-place rolling upgrade, a node's own local
__dbis__seqcursor row (key[Symbol.for('seq'), nodeId], value{seqId, nodes}) can become undecodable —RecordEncoder.decodelogsError decoding record: Data read, but end of buffer not reachedand returnsnull. This is the root cause behind the harper-pro#352 replication wedge; the harper-pro side ships a resilience guard (HarperFast/harper-pro#394) that stops the crash, but the row should not be undecodable in the first place.Evidence (live node, fabric investigation)
Node
8tj-us-west1-a-1after an in-place upgrade to 5.1.1, leaderpq7:systemdatabase__dbis__decoder had its msgpackr shared-structure table fully intact and consistent — 14 structures, on-disk == in-memory, including id 11 =["seqId","nodes"](the correctseqshape).randomAccessStructure = false,typedStructs = 0— structon random-access is not involved.seqentry present,[Symbol(seq), 1](the leader's cursor), decoded tonullwith "end of buffer not reached 0".sendSubscriptionRequestUpdatecrashed on it, inbound replication wedged (connected:true / subscriptions:null / lastReceivedVersion:0), still wedged ~50 min post-boot.So this is not missing structures, not structon typedStructs, and not v4→v5 migration damage (this node was a fresh v5 provision — no LMDB remnants; #1163 / #362 / harper#1275 / copyDb #1260 are all present in 5.1.0/5.1.1 and working). It is a structure-shape skew: the
seqbytes don't cleanly decode against the present, consistent structure table (the decode reads more/fewer fields than the structure it resolves to).Why it doesn't self-heal
hdb_nodesrows self-heal via base-copy resync (independent of the cursor). Theseqcursor only rewrites when replication advances — but the handshake crashes on the undecodable cursor, so replication never advances and the row never rewrites. A self-sustaining wedge (broken on the harper-pro side by harper-pro#394, which lets the node resume and rewrite the row against current structures).Open question / where to look
How does a locally-written
seqrow end up encoded against a structure shape this node later decodes differently across a rolling upgrade? Candidates worth investigating:seqwrite path (core/resources/nodeIdMapping.ts,core/resources/Table.tsseqKeywrites) vs the__dbis__RecordEncoderstructure save/load ordering across the version bump.__dbis__shared-structure table can be appended/reordered such that a row written pre-upgrade references a structure id whose definition differs post-upgrade.Related: harper-pro#352 (the wedge + guard), #1163, #362 (harper#1275), copyDb #1260.
🤖 Filed by Claude Opus 4.8 (1M context), on Kris's behalf.