Repository navigation
Replicated records fail to decode (structon 'end of buffer not reached') — cluster-wide, silent degradation #1348
Description
Activity
Root cause identified (high confidence): shared-structure id divergence across nodes
Digging through the decode path, the failure is a typed-structure (structon) id collision/divergence between the encoder that wrote a record and the decoder reading it — not the migration, and not the documented
0x42metadata-prefix ambiguity.Ruling out the metadata heuristic
The failing buffers start
42 79 ed 6d ba f5 19 33 …. Read as a big-endian float64 that's ≈1.70e12— a valid June-2026 ms timestamp. SoRecordEncoder.decode(resources/RecordEncoder.ts:335) correctly strips the 8-byte RocksDB local-timestamp prefix + metadata flags; the0x42is the timestamp's high byte, not a misread classic record-id #2. The failure is in the value decode that follows (RecordEncoder.ts:404-408,super.decode(buffer.subarray(position, end), end - position)).The mechanism
"Data read, but end of buffer not reached"(msgpackr/unpack.jscheckedRead) means the decoder consumed fewer bytes than the value contains — trailing bytes remain. For a struct/record value this happens when the structure resolved at decode time has a different (shorter) fixed-field layout than the structure used to encode it.Typed-struct ids are assigned locally and sequentially —
recordId = typedStructs.length(structon/struct.js:590and:911). The record header stores that integer id;readStruct(struct.js:952) looks uptypedStructs[recordId]and uses that layout to compute field offsets and total size. There is no content check that the resolved structure matches the bytes. So if node A and node B mint typed structures in different orders, id N denotes a different shape on each node. A record encoded on A against id N, replicated to B, is decoded with B's id-N layout → wrong size → short read → trailing-bytes error. Records can fail on every node (incl. origin) once the structures key — itself a replicated record — is reconciled last-writer-wins to a dictionary whose id N no longer matches what those records were written against.Why none of the existing safeguards catch it
- Reload-on-miss (
structon/index.js:136-153andstruct.js:964-972, the harper#1163 fix) only fires when the structure id is absent (if (!structure)). Here the id is present but the wrong shape, so it never reloads — it silently decodes with the wrong layout. - Save-time CAS
isCompatible(RecordEncoder.ts:285-291,struct.js:1250-1264) only compares dictionary lengths, not the shape at each id. structureVersionpropagated in the audit/replication record (RecordEncoder.ts:796) isstructures.length + typedStructs.length— again a count. Two nodes at the same count can hold divergent shapes.
Structure identity is therefore "position in a locally-grown array + array length," with nothing that detects a same-id/different-shape divergence across nodes.
Likely trigger
Concurrent independent minting on multiple nodes (each assigns the next sequential id to whatever new shape it sees first), then last-writer-wins replication of the
[Symbol.for('structures'), table]key — or a base-copy resync (3-day retention path) overwriting a node's local dictionary after it had already written records against different ids. Either orphans the records written against the losing dictionary. The application'sAppResource.getswallows the throw viaPromise.allSettled+ empty fallback, so it surfaces only as silent loss on ~90 hot keys (see original report).Empirical confirmation (remaining step)
Dump
[Symbol.for('structures'), '<appDb>/<table>']from two nodes' RocksDB and diff the typed-struct entry at the id the failing records reference (extract one full untruncated record via the store and read its header id). Expectation: the same id resolves to different field layouts, and the failing value decodes cleanly against the writer's dictionary.Fix directions (for discussion — this is the distributed shared-structure consistency problem)
- Contain it (smallest): detect the short read at decode (position < end after a struct read), reload structures, and retry; if still mismatched, log once and fall back.
- Make structure identity stable: content-address structure ids (hash of the shape) or namespace typed-struct ids per origin node, so a replicated record never resolves to a divergent shape at the same id.
- Don't LWW the structures key: merge/union structure dictionaries by id with conflict detection rather than last-writer-wins, and carry a per-id shape fingerprint (not just length/count) in
isCompatible/structureVersion.
This touches structon + the harper RecordEncoder structure-replication wiring.
Root-cause analysis compiled by Claude (Opus 4.8) from the structon v1.0.7 source and harper
RecordEncoder.ts(v5.1.3/.4).- Reload-on-miss (
✅ Empirically confirmed — typed-structure dictionaries have forked across nodes
Dumped the live in-memory
typedStructsfor the application's tables on two nodes via CDP (databases.<appDb>.<Table>.primaryStore.encoder.typedStructs) and diffed by id. Each entry below istype:size:keyper field (field names generalized tof0…fn); the byte number is the fixed-section width (what the decoder uses to size the record).Table A— 8 structs on both nodes, divergent at ids[1, 2, 5, 6]:id 1: node A (14B): f0 | f1 | f2(3:1) | f3(3:1) | f4 | f5(16:8) | f6 node D ( 7B): f0 | f1 | f2(3:1) | f3(2:2) | f4 | | f6 id 2: node A ( 8B): f0 | f1 | f2(2:2) | f3(2:2) | f4 | | f6 node D (14B): f0 | f1 | f2(3:1) | f3(3:1) | f4 | f5(16:8) | f6id 1andid 2are essentially swapped between the two nodes, and the fixed-section widths differ (14B vs 7B, 8B vs 14B). A record encoded on node A against id 1 (14-byte fixed section, carries fieldf5) decoded with node D's id-1 layout (7 bytes) consumes 7 fewer bytes →checkedReadfinds leftover bytes →"Data read, but end of buffer not reached". That is precisely the observed error, and the ~65–66 trailing-byte counts are consistent with a missing 8-byte field + width deltas.Other tables diverge the same way:
Table B— 8 structs both nodes, divergent at ids[1, 3, 4, 5, 6, 7](e.g. id 1: node A 5B vs node D 16B).Table C— divergent at ids[1, 2, 3, 4, 5, 6, 7, 9, 10](id 2: node A 2 fields/16B vs node D 6 fields/20B).- Two smaller tables — node A has 1 struct, node D has 0. This is mere absence (reload-on-miss covers it), not divergence — and tellingly these are the tables that don't throw.
Conclusion
The dictionaries genuinely forked: each node minted typed-struct ids in a different order (sequential local assignment,
recordId = typedStructs.length), so the same id denotes a different shape/width per node. Records replicate carrying only the integer id; the receiving node decodes with its own divergent structure at that id → wrong byte width → short read. This matches the symptom on every axis (cluster-wide, same records fail on multiple nodes, only specific records affected — those whose id diverged).Note the fork is symmetric — neither node's dictionary is "correct"; they've diverged from each other. So reload-on-miss cannot fix it (the id is present, just wrong), and there is no single authoritative dictionary to reload from. Recovery requires either decode-time width validation + cross-node structure reconciliation, or re-encoding the affected records under a stable structure identity.
Confirmed via CDP inspection of two live nodes. Dictionaries diffed by id; byte widths computed from
type:sizefield tuples.⚠️ Correction + true root cause (CDP-verified): it's the audit/peer-value decode path, not orphaned stored recordsMy earlier comments framed this as locally-stored records being orphaned by a remapped dictionary. CDP inspection of a live node disproves that. Corrected findings:
1. All persisted records decode cleanly
Scanned the full primary stores via CDP —
Table A(~36k),Table B(~36k),Table C(~36k), and the smaller tables — slicing each value past the metadata prefix and decoding it through the base (msgpackr/structon) decoder:Zero undecodable records. The only decode-to-null cases are legitimate deletion tombstones (value
0xc0= msgpack nil; a few hundred each). So records are correctly re-encoded to local structures on write — locally-stored data is fine. (Divergent dictionaries across nodes are expected/tolerated by design; that's not the bug.)2. The failures are on the bare-value / audit decode path, using local structures
The live error, with the decoded object surfaced, is decisive (field names generalized):
Error: Data read, but end of buffer not reached {"f0":121,"f1":-21,"f2":44,"f3":49,"f4":2198736384,"f5":0} at RecordEncoder.decode (core/resources/RecordEncoder.ts:425:63) ← noMetadata / bare-value branch- The decoded object has the right field names but garbage values (
f0:121,f4:2198736384) — a struct read with the wrong structure: correct keys, wrong byte offsets. Definitive wrong-structure decode (not corruption). - Line 425, not 404 — the
noMetadatabranch. That's the audit-value path:auditStore.ts:579decodes the audit entry's value withstore.decoder.decode(buf, {noMetadata:true})— i.e. the local structures.
3. Why: the audit value is in the origin node's encoding, decoded with local structures
Primary records get re-encoded to local structures on apply (hence clean). But audit entries retain the origin's encoding (the global history is consistent per origin). When a read resolves a value through the audit store — partial/patch reconstruction (
HAS_PARTIAL_RECORD) or out-of-order reconciliation —getValuedecodes that origin-encoded value with the local dictionary. For a peer whose structure ids diverge (measured earlier:Table Aid 1 = 14B on node A vs 7B on node D, id 1/2 swapped), the local structure at that id has the wrong layout → wrong offsets → trailing bytes → caught →null→ empty result.4. The single shared structures key is a cross-node id-space collision point
[Symbol.for('structures'), table]is one id-space per table. harper#1163 (reload-on-miss) and harper-pro#362 (persist replicated structures into that key) handle a missing peer structure, but cannot represent two different shapes at the same id from two origins. So merging peer structures into the shared key can't fix divergence — whichever origin loses the id collision has its audit values misdecode.Fix direction (aligned with the intended "decode-with-peer, re-encode-with-local")
Decode peer-originated values (audit
getValue, and the replication-apply decode) using the origin node's structures, resolved from the audit entry'snodeId+structureVersion(both already carried in the audit record) — instead of the shared local decoder. Keep the local authoritative[structures]key local-only and stable.Update: follow-up investigation refined this further — the actual defect is that the per-origin
HAS_STRUCTURE_UPDATEflag (which drives the receiver to reload structures) is a one-shot signal that can land in a different per-node transaction log than the one whose entries first reference the new structure, so a linear reader misses it. Fixed in #1352 by re-deriving the flag per(log, table)from the monotonicstructureVersion.
CDP-verified on a live node: full primary-store scans (clean), the live error object (wrong-structure garbage), and
auditStore.ts:579/RecordEncoder.ts:425as the decode site.- The decoded object has the right field names but garbage values (
- changed the title
[-]EarlyHints records fail to decode (structon 'end of buffer not reached') — cluster-wide, silent hint degradation[/-][+]Replicated records fail to decode (structon 'end of buffer not reached') — cluster-wide, silent degradation[/+]on Jun 17, 2026 Crash escalation of this same corruption filed as #1370: on the same
eh-prod.gendcluster,httpworkers SIGSEGV with heap-corruption signatures. The hypothesis there is that the native decode path (msgpackr extractor /@harperfast/rocksdb-js) can read out of bounds on these corrupt buffers — which, unlike theRecordEncoder.decodeJStry/catchthat catches thecheckedReadthrow documented here, cannot be caught and crashes the worker. So the same records this issue describes as "silent degradation" can also escalate to a process crash.- added a commit that references this issue
on Jun 18, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
On a production cluster, a subset of application records fail to decode on read with:
The msgpackr decode reads what it believes is a complete value but leaves ~65–66 trailing bytes in the buffer (
checkedRead→ "Data read, but end of buffer not reached"). The value is decoded against a structon typed structure that appears to consume fewer fields than the record actually encodes, so the cursor stops short ofend.Harper version: 5.1.3 (harper-pro). Observed across the cluster.
Scope — cluster-wide, pre-existing, not migration-related
This is not caused by the recent LMDB→RocksDB migration of one node. Every native-RocksDB node (never migrated) exhibits it, at higher rates than the migrated node:
On
node A: 477 errors in 2h across 90 distinct record signatures — i.e. ~90 specific records, each re-failing on every read (13–22 hits each in 2h). The same signatures appear on multiple nodes (e.g.4279ed6e73b26fde…logged on bothnode Aandnode D), confirming the affected records replicate in their broken-to-decode form rather than being node-local corruption.Impact — silent degradation (not an outage)
The triggering path is an application resource (
AppResource.get), which fans out lookups underPromise.allSettledand applies an empty fallback on rejection:So a decode failure does not surface as a client error — the affected field is silently dropped and the response is served incomplete (and the key may be re-enqueued as a cache-miss). Net effect: ~90 popular keys lose the cached optimization on every request and do not self-heal.
Hypothesis
A structon typed-structure desync: the structure resolved at decode time has fewer fields than the structure the record was encoded with, leaving trailing bytes. Candidate mechanisms (need confirmation):
_saveTypedStructures/_ensureTypedStructuresinstructon) not persisted/replicated consistently with the records that reference them, so a reader resolves a stale/base structure.RecordEncoder.ts:404decodesbuffer.subarray(position, end)with explicit lengthend - position; the systematic per-record (not all-records) failure points at record/structure-specific data rather than a global offset bug.Repro material (raw bytes, from logs)
Each log line ends with
data: <hex>(the undecodable value). Samples (note the common42 … 0e0000 40000000 …framing; the trailing string bytes are redacted):Full samples are recoverable from any node's
docker logs(filterError decoding record, the 11th stack line carriesdata:).Suggested next steps
data:payload against the node's stored structon typed structures to confirm the field-count mismatch.Filed from a live production investigation. Compiled by Claude (Opus 4.8) for Kris; figures are from
docker logsat the time of writing.