Skip to content

Replicated records fail to decode (structon 'end of buffer not reached') — cluster-wide, silent degradation #1348

Description

@kriszyp

Summary

On a production cluster, a subset of application records fail to decode on read with:

[error]: Error decoding record Error: Data read, but end of buffer not reached 65
    at checkedRead (msgpackr/unpack.js:217:10)
    at RecordEncoder.unpack (msgpackr/unpack.js:96:12)
    at RecordEncoder.decode (msgpackr/unpack.js:168:15)
    at RecordEncoder.decode (structon/index.js:99:37)
    at <anonymous> (core/resources/RecordEncoder.ts:408:16)
    at decodeFromDatabase (core/resources/blob.ts:968:9)
    at RecordEncoder.decode (core/resources/RecordEncoder.ts:404:37)
    at Store.decodeValue (@harperfast/rocksdb-js/src/store.ts:448:24)
    at <anonymous> (@harperfast/rocksdb-js/src/dbi.ts:196:24)
    at async AppResource.get (.../app-component/src/resources/AppResource.js:137:22)

The msgpackr decode reads what it believes is a complete value but leaves ~65–66 trailing bytes in the buffer (checkedRead → "Data read, but end of buffer not reached"). The value is decoded against a structon typed structure that appears to consume fewer fields than the record actually encodes, so the cursor stops short of end.

Harper version: 5.1.3 (harper-pro). Observed across the cluster.

Scope — cluster-wide, pre-existing, not migration-related

This is not caused by the recent LMDB→RocksDB migration of one node. Every native-RocksDB node (never migrated) exhibits it, at higher rates than the migrated node:

Node "Error decoding record" / 6h
node A 815
node B 493
node C 153
node D (migrated) ~3 / 25 min (lowest)

On node A: 477 errors in 2h across 90 distinct record signatures — i.e. ~90 specific records, each re-failing on every read (13–22 hits each in 2h). The same signatures appear on multiple nodes (e.g. 4279ed6e73b26fde… logged on both node A and node D), confirming the affected records replicate in their broken-to-decode form rather than being node-local corruption.

Impact — silent degradation (not an outage)

The triggering path is an application resource (AppResource.get), which fans out lookups under Promise.allSettled and applies an empty fallback on rejection:

const [aRes, bRes, cRes, dRes] = await Promise.allSettled([...]);
const getValue = (res, fallback) => (res.status === 'fulfilled' ? res.value : fallback);
const partA = getValue(aRes, []);   // decode failure -> []

So a decode failure does not surface as a client error — the affected field is silently dropped and the response is served incomplete (and the key may be re-enqueued as a cache-miss). Net effect: ~90 popular keys lose the cached optimization on every request and do not self-heal.

Hypothesis

A structon typed-structure desync: the structure resolved at decode time has fewer fields than the structure the record was encoded with, leaving trailing bytes. Candidate mechanisms (need confirmation):

  • Typed structures (_saveTypedStructures / _ensureTypedStructures in structon) not persisted/replicated consistently with the records that reference them, so a reader resolves a stale/base structure.
  • A structure evolved (field added) after some records were written, and the older records now resolve against the wrong structure id.

RecordEncoder.ts:404 decodes buffer.subarray(position, end) with explicit length end - position; the systematic per-record (not all-records) failure points at record/structure-specific data rather than a global offset bug.

Repro material (raw bytes, from logs)

Each log line ends with data: <hex> (the undecodable value). Samples (note the common 42 … 0e0000 40000000 … framing; the trailing string bytes are redacted):

4279ed6dbaf519330e0000400000000c42bb <string-bytes…>
4279ed653cd43e890e0000400000000642d937 <string-bytes…>
4279ed6e73b26fde0e0000400000000d41d934 <string-bytes…>  (seen on 2 nodes)

Full samples are recoverable from any node's docker logs (filter Error decoding record, the 11th stack line carries data:).

Suggested next steps

  1. Decode one captured data: payload against the node's stored structon typed structures to confirm the field-count mismatch.
  2. Determine whether the typed structures for these records are present/consistent across nodes (replication of typed structures vs records).
  3. Decide remediation: re-encode the affected records vs a decode-time compatibility path in structon for under-resolved structures.

Filed from a live production investigation. Compiled by Claude (Opus 4.8) for Kris; figures are from docker logs at the time of writing.

Activity

  1. kriszyp commented on Jun 17, 2026

    @kriszyp
    MemberAuthor

    Root cause identified (high confidence): shared-structure id divergence across nodes

    Digging through the decode path, the failure is a typed-structure (structon) id collision/divergence between the encoder that wrote a record and the decoder reading it — not the migration, and not the documented 0x42 metadata-prefix ambiguity.

    Ruling out the metadata heuristic

    The failing buffers start 42 79 ed 6d ba f5 19 33 …. Read as a big-endian float64 that's ≈ 1.70e12 — a valid June-2026 ms timestamp. So RecordEncoder.decode (resources/RecordEncoder.ts:335) correctly strips the 8-byte RocksDB local-timestamp prefix + metadata flags; the 0x42 is the timestamp's high byte, not a misread classic record-id #2. The failure is in the value decode that follows (RecordEncoder.ts:404-408, super.decode(buffer.subarray(position, end), end - position)).

    The mechanism

    "Data read, but end of buffer not reached" (msgpackr/unpack.js checkedRead) means the decoder consumed fewer bytes than the value contains — trailing bytes remain. For a struct/record value this happens when the structure resolved at decode time has a different (shorter) fixed-field layout than the structure used to encode it.

    Typed-struct ids are assigned locally and sequentially — recordId = typedStructs.length (structon/struct.js:590 and :911). The record header stores that integer id; readStruct (struct.js:952) looks up typedStructs[recordId] and uses that layout to compute field offsets and total size. There is no content check that the resolved structure matches the bytes. So if node A and node B mint typed structures in different orders, id N denotes a different shape on each node. A record encoded on A against id N, replicated to B, is decoded with B's id-N layout → wrong size → short read → trailing-bytes error. Records can fail on every node (incl. origin) once the structures key — itself a replicated record — is reconciled last-writer-wins to a dictionary whose id N no longer matches what those records were written against.

    Why none of the existing safeguards catch it

    1. Reload-on-miss (structon/index.js:136-153 and struct.js:964-972, the harper#1163 fix) only fires when the structure id is absent (if (!structure)). Here the id is present but the wrong shape, so it never reloads — it silently decodes with the wrong layout.
    2. Save-time CAS isCompatible (RecordEncoder.ts:285-291, struct.js:1250-1264) only compares dictionary lengths, not the shape at each id.
    3. structureVersion propagated in the audit/replication record (RecordEncoder.ts:796) is structures.length + typedStructs.length — again a count. Two nodes at the same count can hold divergent shapes.

    Structure identity is therefore "position in a locally-grown array + array length," with nothing that detects a same-id/different-shape divergence across nodes.

    Likely trigger

    Concurrent independent minting on multiple nodes (each assigns the next sequential id to whatever new shape it sees first), then last-writer-wins replication of the [Symbol.for('structures'), table] key — or a base-copy resync (3-day retention path) overwriting a node's local dictionary after it had already written records against different ids. Either orphans the records written against the losing dictionary. The application's AppResource.get swallows the throw via Promise.allSettled + empty fallback, so it surfaces only as silent loss on ~90 hot keys (see original report).

    Empirical confirmation (remaining step)

    Dump [Symbol.for('structures'), '<appDb>/<table>'] from two nodes' RocksDB and diff the typed-struct entry at the id the failing records reference (extract one full untruncated record via the store and read its header id). Expectation: the same id resolves to different field layouts, and the failing value decodes cleanly against the writer's dictionary.

    Fix directions (for discussion — this is the distributed shared-structure consistency problem)

    • Contain it (smallest): detect the short read at decode (position < end after a struct read), reload structures, and retry; if still mismatched, log once and fall back.
    • Make structure identity stable: content-address structure ids (hash of the shape) or namespace typed-struct ids per origin node, so a replicated record never resolves to a divergent shape at the same id.
    • Don't LWW the structures key: merge/union structure dictionaries by id with conflict detection rather than last-writer-wins, and carry a per-id shape fingerprint (not just length/count) in isCompatible / structureVersion.

    This touches structon + the harper RecordEncoder structure-replication wiring.


    Root-cause analysis compiled by Claude (Opus 4.8) from the structon v1.0.7 source and harper RecordEncoder.ts (v5.1.3/.4).

  2. kriszyp commented on Jun 17, 2026

    @kriszyp
    MemberAuthor

    ✅ Empirically confirmed — typed-structure dictionaries have forked across nodes

    Dumped the live in-memory typedStructs for the application's tables on two nodes via CDP (databases.<appDb>.<Table>.primaryStore.encoder.typedStructs) and diffed by id. Each entry below is type:size:key per field (field names generalized to f0…fn); the byte number is the fixed-section width (what the decoder uses to size the record).

    Table A — 8 structs on both nodes, divergent at ids [1, 2, 5, 6]:

    id 1:  node A (14B): f0 | f1 | f2(3:1) | f3(3:1) | f4 | f5(16:8) | f6
           node D ( 7B): f0 | f1 | f2(3:1) | f3(2:2) | f4 |           | f6
    id 2:  node A ( 8B): f0 | f1 | f2(2:2) | f3(2:2) | f4 |           | f6
           node D (14B): f0 | f1 | f2(3:1) | f3(3:1) | f4 | f5(16:8) | f6
    

    id 1 and id 2 are essentially swapped between the two nodes, and the fixed-section widths differ (14B vs 7B, 8B vs 14B). A record encoded on node A against id 1 (14-byte fixed section, carries field f5) decoded with node D's id-1 layout (7 bytes) consumes 7 fewer bytes → checkedRead finds leftover bytes → "Data read, but end of buffer not reached". That is precisely the observed error, and the ~65–66 trailing-byte counts are consistent with a missing 8-byte field + width deltas.

    Other tables diverge the same way:

    • Table B — 8 structs both nodes, divergent at ids [1, 3, 4, 5, 6, 7] (e.g. id 1: node A 5B vs node D 16B).
    • Table C — divergent at ids [1, 2, 3, 4, 5, 6, 7, 9, 10] (id 2: node A 2 fields/16B vs node D 6 fields/20B).
    • Two smaller tables — node A has 1 struct, node D has 0. This is mere absence (reload-on-miss covers it), not divergence — and tellingly these are the tables that don't throw.

    Conclusion

    The dictionaries genuinely forked: each node minted typed-struct ids in a different order (sequential local assignment, recordId = typedStructs.length), so the same id denotes a different shape/width per node. Records replicate carrying only the integer id; the receiving node decodes with its own divergent structure at that id → wrong byte width → short read. This matches the symptom on every axis (cluster-wide, same records fail on multiple nodes, only specific records affected — those whose id diverged).

    Note the fork is symmetric — neither node's dictionary is "correct"; they've diverged from each other. So reload-on-miss cannot fix it (the id is present, just wrong), and there is no single authoritative dictionary to reload from. Recovery requires either decode-time width validation + cross-node structure reconciliation, or re-encoding the affected records under a stable structure identity.


    Confirmed via CDP inspection of two live nodes. Dictionaries diffed by id; byte widths computed from type:size field tuples.

  3. kriszyp commented on Jun 17, 2026

    @kriszyp
    MemberAuthor

    ⚠️ Correction + true root cause (CDP-verified): it's the audit/peer-value decode path, not orphaned stored records

    My earlier comments framed this as locally-stored records being orphaned by a remapped dictionary. CDP inspection of a live node disproves that. Corrected findings:

    1. All persisted records decode cleanly

    Scanned the full primary stores via CDP — Table A (~36k), Table B (~36k), Table C (~36k), and the smaller tables — slicing each value past the metadata prefix and decoding it through the base (msgpackr/structon) decoder:

    Zero undecodable records. The only decode-to-null cases are legitimate deletion tombstones (value 0xc0 = msgpack nil; a few hundred each). So records are correctly re-encoded to local structures on write — locally-stored data is fine. (Divergent dictionaries across nodes are expected/tolerated by design; that's not the bug.)

    2. The failures are on the bare-value / audit decode path, using local structures

    The live error, with the decoded object surfaced, is decisive (field names generalized):

    Error: Data read, but end of buffer not reached
      {"f0":121,"f1":-21,"f2":44,"f3":49,"f4":2198736384,"f5":0}
      at RecordEncoder.decode (core/resources/RecordEncoder.ts:425:63)   ← noMetadata / bare-value branch
    
    • The decoded object has the right field names but garbage values (f0:121, f4:2198736384) — a struct read with the wrong structure: correct keys, wrong byte offsets. Definitive wrong-structure decode (not corruption).
    • Line 425, not 404 — the noMetadata branch. That's the audit-value path: auditStore.ts:579 decodes the audit entry's value with store.decoder.decode(buf, {noMetadata:true}) — i.e. the local structures.

    3. Why: the audit value is in the origin node's encoding, decoded with local structures

    Primary records get re-encoded to local structures on apply (hence clean). But audit entries retain the origin's encoding (the global history is consistent per origin). When a read resolves a value through the audit store — partial/patch reconstruction (HAS_PARTIAL_RECORD) or out-of-order reconciliation — getValue decodes that origin-encoded value with the local dictionary. For a peer whose structure ids diverge (measured earlier: Table A id 1 = 14B on node A vs 7B on node D, id 1/2 swapped), the local structure at that id has the wrong layout → wrong offsets → trailing bytes → caught → null → empty result.

    4. The single shared structures key is a cross-node id-space collision point

    [Symbol.for('structures'), table] is one id-space per table. harper#1163 (reload-on-miss) and harper-pro#362 (persist replicated structures into that key) handle a missing peer structure, but cannot represent two different shapes at the same id from two origins. So merging peer structures into the shared key can't fix divergence — whichever origin loses the id collision has its audit values misdecode.

    Fix direction (aligned with the intended "decode-with-peer, re-encode-with-local")

    Decode peer-originated values (audit getValue, and the replication-apply decode) using the origin node's structures, resolved from the audit entry's nodeId + structureVersion (both already carried in the audit record) — instead of the shared local decoder. Keep the local authoritative [structures] key local-only and stable.

    Update: follow-up investigation refined this further — the actual defect is that the per-origin HAS_STRUCTURE_UPDATE flag (which drives the receiver to reload structures) is a one-shot signal that can land in a different per-node transaction log than the one whose entries first reference the new structure, so a linear reader misses it. Fixed in #1352 by re-deriving the flag per (log, table) from the monotonic structureVersion.


    CDP-verified on a live node: full primary-store scans (clean), the live error object (wrong-structure garbage), and auditStore.ts:579 / RecordEncoder.ts:425 as the decode site.

  4. changed the title [-]EarlyHints records fail to decode (structon 'end of buffer not reached') — cluster-wide, silent hint degradation[/-] [+]Replicated records fail to decode (structon 'end of buffer not reached') — cluster-wide, silent degradation[/+] on Jun 17, 2026
  5. maurice-harper commented on Jun 18, 2026

    @maurice-harper
    Contributor

    Crash escalation of this same corruption filed as #1370: on the same eh-prod.gend cluster, http workers SIGSEGV with heap-corruption signatures. The hypothesis there is that the native decode path (msgpackr extractor / @harperfast/rocksdb-js) can read out of bounds on these corrupt buffers — which, unlike the RecordEncoder.decode JS try/catch that catches the checkedRead throw documented here, cannot be caught and crashes the worker. So the same records this issue describes as "silent degradation" can also escalate to a process crash.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions