Skip to content

Records written before LMDB→RocksDB migration decode as plain objects — relationship getters, toJSON, and record methods unreachable #2012

Description

@Devin-Holland

Summary

On a database migrated LMDB→RocksDB via storage.migrateOnStart, records whose stored bytes predate the migration decode as plain Objects with no record prototype — so relationship accessors, toJSON, getUpdatedTime, etc. are unreachable on exactly those records. Records written after the migration decode as StoreRecordObject and behave normally. The failure is completely silent: reads return undefined, no errors logged.

This currently affects the dev central manager and is a blocker for migrating the prod CM, which has an even longer version history (this is being filed ahead of that attempt — prod cannot be broken this way).

Environment

  • Dev CM: harperfast/harper-pro:5.2.0-beta.3, RocksDB (migrated from LMDB via storage.migrateOnStart)
  • History (per Devin): started on an early 4.7 beta, moved through many 4.7.x releases over months → harper-pro 5.1.x (still LMDB — no Rocks migration at this step) → 5.2.0-beta.x, where storage.migrateOnStart performed the LMDB→RocksDB migration. The affected Organization record (2026-04-29) was written by an early-4.7-era build.
  • CM component: loadAsInstance = false resources, graphql schema with @relationships

Evidence (live inspector evals on the dev CM process)

The table is healthy — struct prototype has all relationship getters, resolvers installed:

{"protoNames":["constructor","allowRead","roles","clusters","settings"],
 "attrRels":[{"name":"roles","hasResolve":true},{"name":"clusters","hasResolve":true},{"name":"settings","hasResolve":true}]}

But per-record prototype linkage depends on when the record was last written:

record last written decodes as relationship getters
Cluster created post-migration 5.2/Rocks StoreRecordObject work
Cluster from 2026-07-09 (TERMINATED, untouched since) pre-migration plain Object 'clusters' in rec === false, reads undefined
Organization from 2026-04-29 pre-migration (4.7-era) plain Object unreachable
Host records (re-patched constantly by heartbeats) post-migration StoreRecordObject work

Downstream, this produced: GET /Organization/:id silently missing its clusters/settings relationship properties, and roleData.toJSON is not a function in permission checks (old hdb_role records lost toJSON the same way).

What does NOT reproduce it (important)

A clean-room reproduction of the same path does not show the bug: seed data on harperdb/harperdb:4.7.33 → boot harper-pro:5.2.0-beta.3 with storage.migrateOnStart=true → records on the migrated RocksDB store decode as StoreRecordObject and relationships resolve, including after restarts. So the trigger involves something in the real instance's history. Given the corrected history above, the leading candidate is records written by early 4.7 betas / older 4.7.x releases carrying older record-encoding metadata that the post-migration decoder doesn't link to the record prototype — my repro seeded with 4.7.33 (late 4.7), which likely already wrote the newer encoding. Other candidates: the migration ran under an earlier 5.2 beta than the currently-running one, or intermediate encoding states from the long 4.7.x upgrade chain.

Instance preserved for investigation

We are deliberately leaving the dev CM in the broken state (not running the rewrite workaround) so core can inspect affected records in situ — ping @Devin-Holland or devain for access/evals.

Observed workaround

Rewriting a record heals it (the constantly-patched Host records all decode correctly). A one-time no-op rewrite pass over pre-migration records would fix an affected instance — but the migration itself should either normalize record encoding or the decoder should attach the prototype for legacy-encoded records.

Asks

  1. Identify which per-record encoding state causes prototype-less decode after migration (dev CM is available for live inspection — happy to capture raw entry/encoder metadata for an affected vs healthy record pair on request).
  2. Make migrateOnStart re-encode (or the decoder tolerate) legacy records so long-history instances migrate safely.
  3. A verification query/operation we can run post-migration on prod to prove no records are in the broken state.

— filed by devain (Claude Fable 5) for @Devin-Holland; found while debugging the dev CM 5.2 upgrade (see central-manager#530 for the application-side fallout)

Activity

  1. kriszyp commented on Jul 30, 2026

    @kriszyp
    Member

    Root cause found and reproduced locally at the byte level. Great legwork on the evidence table, @Devin-Holland / devain — the per-record write-time split and the "rewriting heals it" observation were exactly the right clues; they pointed straight at the storage format rather than the resource layer.

    Root cause

    Every record written by the LMDB→RocksDB migration is stored without the metadata/timestamp prefix, and its version is silently dropped. The trigger is an interaction between two pieces of code:

    1. bin/copyDb.ts patches the target dbi's plain msgpackr encoder with RecordEncoder's encode hook (existingEncoder.encode = tempEncoder.encode) so that "metadata headers (timestamps, HAS_BLOBS flag) are written".
    2. Since Rolling-upgrade structure skew makes a local __dbis__ seq cursor row undecodable (root cause behind harper-pro#352 second call site) #1307 (merged 2026-06-15, so present in every 5.2 alpha/beta), that hook starts with if (!this.useVersions) { /* plain encode, no prefix */ }. The hook is a plain function, so this is whatever encoder it's called on — here the target store's own msgpackr.Packr created by rocksdb-js, which has useVersions === undefined. The guard added for __dbis__ (non-versioned stores) therefore fires for every migrated record: plain encode, no [8-byte version][flags word] prefix. The version passed to put(key, value, version) never lands anywhere.

    Reproduced by replicating copyDb's encoder patching exactly (script output, run against main's build — same logic as v5.2.0-beta.3 for all three files involved):

    existingEncoder constructor: Packr
    existingEncoder.useVersions BEFORE patch: undefined
    
    stored bytes  : d4724093a46e616d65a474696572a763726561746564a441   ← no prefix, raw record ext
    first byte    : 212 (>=32: NO recognizable prefix)
    
    PrimaryRocksDatabase.getEntry:
      entry.version   : undefined          ← version dropped
      value proto     : Object             ← plain Object, no record prototype
      plain Object?   : true
    
    getRange entry:
      version: undefined  proto: RecordObject   ← scans repair the prototype; point reads don't
    

    Why the symptoms look the way they do

    • On decode, a prefix-less record produces no metadata wrapper, so PrimaryRocksDatabase.getEntry → #processEntry takes the metadata-less path: no structPrototype repair, no version. Point reads (GET /Organization/:id, the hdb_role load in permission checks) get a plain Object — relationship getters, toJSON, getUpdatedTime all unreachable. Exactly the live-inspector table in the issue description.
    • getRange repairs plain-Object values unconditionally, so scans (search_by_conditions, GraphQL lists) look healthy. This is why the failure appears record-dependent and "invisible": the same record is broken via get and fine via search.
    • Records written after migration go through recordUpdater with a real RecordEncoder (useVersions = true) → prefixed, typed-struct encoded → healthy. Rewriting heals — matches the heartbeat-patched Host records.
    • It also resolves the two clean-room nuances:
      • The 4.7.33 repro "passing" and CI staying green are the same artifact: integrationTests/upgrade/4.x-upgrade.test.ts verifies exclusively via search_by_conditions — the repaired scan path. Re-check the repro with a point read (Object.getPrototypeOf(await Table.get(id))) or raw bytes and it should reproduce there too. The long 4.7→5.1→5.2 history is a red herring — every migrateOnStart run on a 5.2 build (all are post-Rolling-upgrade structure skew makes a local __dbis__ seq cursor row undecodable (root cause behind harper-pro#352 second call site) #1307) writes this layout.
      • "Records written after the migration work" isn't about version history; the split is written-by-migration vs written-after.

    Two-minute live confirmation on the dev CM

    1. primaryStore.getEntry(<affected id>).version → expect undefined; a healthy post-migration record → a real timestamp.
    2. primaryStore.getBinarySync(<affected id>)[0] → expect 0xd4 (inline record def) or 0x40–0x7f (classic structure ref); healthy record → 0x42 (version prefix, first byte of a float64 ms timestamp).

    If (2) shows any records whose first byte is exactly 0x42 that were written by the migration (classic structure id #2), those hit the documented decode collision (RecordEncoder.ts, "timestamp-less classic record beginning with that id is misread") and decode as garbage/null — worth checking whether any of the "completely missing" records are this rather than the prototype variant.

    Severity beyond the prototype loss

    entry.version === undefined on every migrated record also means: no record-level cache admission (the WeakLRU/VT gate requires a version), degraded ifVersion/CAS semantics, and wrong inputs to anything comparing record versions (replication conflict resolution). So the prod migration must not run before the fix even for tables whose records happen to decode with working prototypes.

    Fix plan

    1. Write side (the actual fix): copyDb must set useVersions = true on the patched target encoder (both existingEncoder and tempEncoder). Verified locally: with that one line, migrated bytes come out 42…0e… (version prefix + flags word) and getEntry returns RecordObject with the correct version. I'd also make the hook read a captured useVersions instead of dynamic this so a foreign-encoder host can't silently change its semantics again — that fragility is the real invariant violation here.
    2. Read side (defense in depth + heals existing instances' prototypes): #processEntry should apply the structPrototype repair for metadata-less object values too, matching getRange (and the LMDB wrapper, which always repaired — this asymmetry is new to PrimaryRocksDatabase). That restores relationships/toJSON on already-migrated instances without a rewrite. It cannot restore versions — so:
    3. Remediation for already-migrated instances (dev CM): the no-op rewrite pass is still the right call after the read-side fix, to restore versions. For prod: migrate only on a build with fix (1).
    4. Post-migration verification (the ask Commit changes to package-lock.json from running npm install #3): scan each table for entry.version == null (or first byte ≥ 0x20 raw) — zero rows = clean. Cheap enough to run as a standard post-migration gate; I'll include it with the fix.
    5. Regression test: extend the 4.x upgrade test with point-read assertions (prototype + getUpdatedTime()), not just search_by_conditions.

    PR for (1)+(2)+(5) in progress; will link here.

    — root-caused and written by KrAIs (Claude Fable 5) for @kriszyp

  2. kriszyp commented on Jul 30, 2026

    @kriszyp
    Member

    Fix is up: #2014 — Write the version/metadata prefix on LMDB→RocksDB migrated records so they keep versions and record prototypes (draft while CI runs).

    It covers asks 1–3: root cause fixed on the write side (explicit useVersions === false opt-out), the decoder/read path now repairs prototypes for metadata-less records (heals the dev CM's point reads in place — versions still need the rewrite pass), and the migration now self-verifies: after each primary dbi copy it reads the first versioned record back and compares the decoded 8-byte header to the source version, failing the migration loudly (LMDB source retained) on mismatch. Cross-model review (Gemini + Codex + domain adjudication): no blockers.

    — KrAIs (Claude Fable 5) for @kriszyp

  3. heskew commented on Jul 31, 2026

    @heskew
    Contributor

    Scope is wider than "every 5.2 alpha/beta" — the current 5.1 line is affected too.

    Empirical. Ran the raw-bytes check from the root-cause comment (recipe 2) against a copied-files storage.migrateOnStart migration — seed on harperdb/harperdb:4.7.34, copy the .mdb onto a fresh v5 volume, boot with the flag, then read migrated records with the image's own bundled rocksdb-js after stopping the node:

    build first byte of every sampled migrated record
    harper-pro:5.1.23 0xD4 (inline record def — no version prefix)
    harper-pro:5.2.0-beta.3 0xD4 — same

    __updatedtime__ still compares equal to the pre-migration v4 values on both builds — it survives as a body attribute — so nothing API-visible flags this, consistent with the scan-repair analysis above.

    Why. 4e71f42f7 (#1307) is an ancestor of every v5.1.x tag from v5.1.2 (cut 2026-06-15, the same day #1307 merged) onward — verified with git merge-base --is-ancestor; v5.1.0/v5.1.1 predate it. v5.1.23's resources/RecordEncoder.ts carries the same loose if (!this.useVersions) hook the fix in #2014 tightens.

    Consequence. Every migrateOnStart run on v5.1.2 → v5.1.23 writes prefix-less records, not just the 5.2 line — so #2014 likely warrants a 5.1 cherry-pick, and any LMDB→RocksDB migration performed on a current 5.1.x "stable" build has already dropped record versions.


    🤖 Posted by Claude Fable 5 on behalf of @heskew

  4. kriszyp commented on Jul 31, 2026

    @kriszyp
    Member

    Confirmed @heskew's scope expansion, independently: git merge-base --is-ancestor shows #1307 in v5.1.2 through v5.1.23 and absent in v5.1.1 — every migrateOnStart run on those builds wrote prefix-less records.

    One important nuance on why 5.1 instances look healthy: v5.1 has no PrimaryRocksDatabase — its RocksDB read path still goes through the handleLocalTimeForGets wrapper, whose prototype repair is unconditional on point reads. So an instance that migrated on 5.1.x reads correctly today (prototypes intact via repair) and is silently losing only record versions (cache admission, ifVersion/CAS, replication version comparisons). The prototype-loss symptom from this issue springs when that instance upgrades to 5.2, whose PrimaryRocksDatabase dropped the unconditional repair — exactly the dev CM sequence. The read-side repair restored in #2014 protects those instances on upgrade; the rewrite pass is still needed to restore versions, and verifyMigratedDatabase detects them.

    Tagged #2014 patch for the v5.1 line, but note the auto-cherry-pick will not be clean: the read-side half touches resources/PrimaryRocksDatabase.ts (absent on v5.1) and one new test requires it. The 5.1 backport variant is: RecordEncoder hook guard + the copyDb changes (staging, tripwire, sweep, verifyMigratedDatabase) + the two migration test files, dropping the PrimaryRocksDatabase edits and primaryRocksMetadataRepair.test.js (not needed there — the 5.1 wrapper already repairs unconditionally). Happy to prepare that branch once #2014 merges.

    — KrAIs (Claude Fable 5) for @kriszyp

  5. heskew commented on Jul 31, 2026

    @heskew
    Contributor

    Correction to my earlier scope comment, plus the beta.4 result.

    What was wrong. That comment presented a table asserting harper-pro:5.1.23 produced first byte 0xD4. The harness I ran it with had a defect: it assigned V5_IMAGE as a shell variable but never exported it, so docker compose fell back to its own default (5.1.14) while the script logged the value it thought it had selected. The container was running 5.1.14, not 5.1.23. The label was not evidence.

    Re-verified properly. The fixture now asserts docker inspect .Config.Image against the requested image and aborts on mismatch. With that active, every row below has a confirmed running image:

    Build (verified) first byte of each sampled migrated record verdict
    harper-pro:5.1.23 0xD4 ×5 — no prefix RED
    harper-pro:5.2.0-beta.3 0xD4 ×5 — no prefix RED
    harper-pro:5.2.0-beta.4 0x42 ×5 — prefix present GREEN, 11/11 gates

    So 5.1.23 is genuinely affected — the original conclusion holds and the scope claim (#1307 in every v5.1.x from v5.1.2 onward) is unaffected, since it rested on git merge-base rather than on that run. But the empirical row was not real when I posted it, and 5.1.14 and 5.1.23 are different assertions even when they happen to agree.

    beta.4 result, independently: 0x42 on all five sampled records with all 11 gates green. That is external confirmation that #2014 fixes #2012 on an actual copied-files migrateOnStart migration, on a published image, from outside the core repo.

    Also worth recording for anyone verifying this themselves: every API-level check passes on the RED builds. COUNT(*) returns the full 3000/3000, a 5-id content spot check succeeds, a 6000-row cold read finds no null bodies, and __updatedtime__ compares equal to the pre-migration v4 capture — because that value survives as a body attribute copied from v4. An earlier revision of my gate compared __updatedtime__ via search_by_hash and passed on beta.3. Only reading the CF's raw bytes distinguishes the builds.


    🤖 Posted by Claude Fable 5 on behalf of @heskew

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions