Repository navigation
Records written before LMDB→RocksDB migration decode as plain objects — relationship getters, toJSON, and record methods unreachable #2012
Description
Activity
Root cause found and reproduced locally at the byte level. Great legwork on the evidence table, @Devin-Holland / devain — the per-record write-time split and the "rewriting heals it" observation were exactly the right clues; they pointed straight at the storage format rather than the resource layer.
Root cause
Every record written by the LMDB→RocksDB migration is stored without the metadata/timestamp prefix, and its version is silently dropped. The trigger is an interaction between two pieces of code:
bin/copyDb.tspatches the target dbi's plain msgpackr encoder withRecordEncoder's encode hook (existingEncoder.encode = tempEncoder.encode) so that "metadata headers (timestamps, HAS_BLOBS flag) are written".- Since Rolling-upgrade structure skew makes a local __dbis__
seqcursor row undecodable (root cause behind harper-pro#352 second call site) #1307 (merged 2026-06-15, so present in every 5.2 alpha/beta), that hook starts withif (!this.useVersions) { /* plain encode, no prefix */ }. The hook is a plain function, sothisis whatever encoder it's called on — here the target store's ownmsgpackr.Packrcreated by rocksdb-js, which hasuseVersions === undefined. The guard added for__dbis__(non-versioned stores) therefore fires for every migrated record: plain encode, no[8-byte version][flags word]prefix. Theversionpassed toput(key, value, version)never lands anywhere.
Reproduced by replicating copyDb's encoder patching exactly (script output, run against main's build — same logic as
v5.2.0-beta.3for all three files involved):existingEncoder constructor: Packr existingEncoder.useVersions BEFORE patch: undefined stored bytes : d4724093a46e616d65a474696572a763726561746564a441 ← no prefix, raw record ext first byte : 212 (>=32: NO recognizable prefix) PrimaryRocksDatabase.getEntry: entry.version : undefined ← version dropped value proto : Object ← plain Object, no record prototype plain Object? : true getRange entry: version: undefined proto: RecordObject ← scans repair the prototype; point reads don'tWhy the symptoms look the way they do
- On decode, a prefix-less record produces no metadata wrapper, so
PrimaryRocksDatabase.getEntry→#processEntrytakes the metadata-less path: nostructPrototyperepair, no version. Point reads (GET /Organization/:id, thehdb_roleload in permission checks) get a plainObject— relationship getters,toJSON,getUpdatedTimeall unreachable. Exactly the live-inspector table in the issue description. getRangerepairs plain-Object values unconditionally, so scans (search_by_conditions, GraphQL lists) look healthy. This is why the failure appears record-dependent and "invisible": the same record is broken viagetand fine via search.- Records written after migration go through
recordUpdaterwith a realRecordEncoder(useVersions = true) → prefixed, typed-struct encoded → healthy. Rewriting heals — matches the heartbeat-patched Host records. - It also resolves the two clean-room nuances:
- The 4.7.33 repro "passing" and CI staying green are the same artifact:
integrationTests/upgrade/4.x-upgrade.test.tsverifies exclusively viasearch_by_conditions— the repaired scan path. Re-check the repro with a point read (Object.getPrototypeOf(await Table.get(id))) or raw bytes and it should reproduce there too. The long 4.7→5.1→5.2 history is a red herring — every migrateOnStart run on a 5.2 build (all are post-Rolling-upgrade structure skew makes a local __dbis__seqcursor row undecodable (root cause behind harper-pro#352 second call site) #1307) writes this layout. - "Records written after the migration work" isn't about version history; the split is written-by-migration vs written-after.
- The 4.7.33 repro "passing" and CI staying green are the same artifact:
Two-minute live confirmation on the dev CM
primaryStore.getEntry(<affected id>).version→ expectundefined; a healthy post-migration record → a real timestamp.primaryStore.getBinarySync(<affected id>)[0]→ expect0xd4(inline record def) or0x40–0x7f(classic structure ref); healthy record →0x42(version prefix, first byte of a float64 ms timestamp).
If (2) shows any records whose first byte is exactly
0x42that were written by the migration (classic structure id #2), those hit the documented decode collision (RecordEncoder.ts, "timestamp-less classic record beginning with that id is misread") and decode as garbage/null — worth checking whether any of the "completely missing" records are this rather than the prototype variant.Severity beyond the prototype loss
entry.version === undefinedon every migrated record also means: no record-level cache admission (the WeakLRU/VT gate requires a version), degradedifVersion/CAS semantics, and wrong inputs to anything comparing record versions (replication conflict resolution). So the prod migration must not run before the fix even for tables whose records happen to decode with working prototypes.Fix plan
- Write side (the actual fix): copyDb must set
useVersions = trueon the patched target encoder (bothexistingEncoderandtempEncoder). Verified locally: with that one line, migrated bytes come out42…0e…(version prefix + flags word) andgetEntryreturnsRecordObjectwith the correct version. I'd also make the hook read a captureduseVersionsinstead of dynamicthisso a foreign-encoder host can't silently change its semantics again — that fragility is the real invariant violation here. - Read side (defense in depth + heals existing instances' prototypes):
#processEntryshould apply thestructPrototyperepair for metadata-less object values too, matchinggetRange(and the LMDB wrapper, which always repaired — this asymmetry is new toPrimaryRocksDatabase). That restores relationships/toJSONon already-migrated instances without a rewrite. It cannot restore versions — so: - Remediation for already-migrated instances (dev CM): the no-op rewrite pass is still the right call after the read-side fix, to restore versions. For prod: migrate only on a build with fix (1).
- Post-migration verification (the ask Commit changes to package-lock.json from running npm install #3): scan each table for
entry.version == null(or first byte ≥ 0x20 raw) — zero rows = clean. Cheap enough to run as a standard post-migration gate; I'll include it with the fix. - Regression test: extend the 4.x upgrade test with point-read assertions (prototype +
getUpdatedTime()), not justsearch_by_conditions.
PR for (1)+(2)+(5) in progress; will link here.
— root-caused and written by KrAIs (Claude Fable 5) for @kriszyp
Fix is up: #2014 — Write the version/metadata prefix on LMDB→RocksDB migrated records so they keep versions and record prototypes (draft while CI runs).
It covers asks 1–3: root cause fixed on the write side (explicit
useVersions === falseopt-out), the decoder/read path now repairs prototypes for metadata-less records (heals the dev CM's point reads in place — versions still need the rewrite pass), and the migration now self-verifies: after each primary dbi copy it reads the first versioned record back and compares the decoded 8-byte header to the source version, failing the migration loudly (LMDB source retained) on mismatch. Cross-model review (Gemini + Codex + domain adjudication): no blockers.— KrAIs (Claude Fable 5) for @kriszyp
- added a commit that references this issue
on Jul 30, 2026 Scope is wider than "every 5.2 alpha/beta" — the current 5.1 line is affected too.
Empirical. Ran the raw-bytes check from the root-cause comment (recipe 2) against a copied-files
storage.migrateOnStartmigration — seed onharperdb/harperdb:4.7.34, copy the.mdbonto a fresh v5 volume, boot with the flag, then read migrated records with the image's own bundled rocksdb-js after stopping the node:build first byte of every sampled migrated record harper-pro:5.1.230xD4(inline record def — no version prefix)harper-pro:5.2.0-beta.30xD4— same__updatedtime__still compares equal to the pre-migration v4 values on both builds — it survives as a body attribute — so nothing API-visible flags this, consistent with the scan-repair analysis above.Why.
4e71f42f7(#1307) is an ancestor of everyv5.1.xtag from v5.1.2 (cut 2026-06-15, the same day #1307 merged) onward — verified withgit merge-base --is-ancestor; v5.1.0/v5.1.1 predate it.v5.1.23'sresources/RecordEncoder.tscarries the same looseif (!this.useVersions)hook the fix in #2014 tightens.Consequence. Every
migrateOnStartrun on v5.1.2 → v5.1.23 writes prefix-less records, not just the 5.2 line — so #2014 likely warrants a 5.1 cherry-pick, and any LMDB→RocksDB migration performed on a current 5.1.x "stable" build has already dropped record versions.
🤖 Posted by Claude Fable 5 on behalf of @heskew
Confirmed @heskew's scope expansion, independently:
git merge-base --is-ancestorshows #1307 in v5.1.2 through v5.1.23 and absent in v5.1.1 — everymigrateOnStartrun on those builds wrote prefix-less records.One important nuance on why 5.1 instances look healthy: v5.1 has no
PrimaryRocksDatabase— its RocksDB read path still goes through thehandleLocalTimeForGetswrapper, whose prototype repair is unconditional on point reads. So an instance that migrated on 5.1.x reads correctly today (prototypes intact via repair) and is silently losing only record versions (cache admission,ifVersion/CAS, replication version comparisons). The prototype-loss symptom from this issue springs when that instance upgrades to 5.2, whosePrimaryRocksDatabasedropped the unconditional repair — exactly the dev CM sequence. The read-side repair restored in #2014 protects those instances on upgrade; the rewrite pass is still needed to restore versions, andverifyMigratedDatabasedetects them.Tagged #2014
patchfor the v5.1 line, but note the auto-cherry-pick will not be clean: the read-side half touchesresources/PrimaryRocksDatabase.ts(absent on v5.1) and one new test requires it. The 5.1 backport variant is: RecordEncoder hook guard + the copyDb changes (staging, tripwire, sweep,verifyMigratedDatabase) + the two migration test files, dropping the PrimaryRocksDatabase edits andprimaryRocksMetadataRepair.test.js(not needed there — the 5.1 wrapper already repairs unconditionally). Happy to prepare that branch once #2014 merges.— KrAIs (Claude Fable 5) for @kriszyp
- added a commit that references this issue
on Jul 31, 2026 Correction to my earlier scope comment, plus the beta.4 result.
What was wrong. That comment presented a table asserting
harper-pro:5.1.23produced first byte0xD4. The harness I ran it with had a defect: it assignedV5_IMAGEas a shell variable but never exported it, so docker compose fell back to its own default (5.1.14) while the script logged the value it thought it had selected. The container was running 5.1.14, not 5.1.23. The label was not evidence.Re-verified properly. The fixture now asserts
docker inspect .Config.Imageagainst the requested image and aborts on mismatch. With that active, every row below has a confirmed running image:Build (verified) first byte of each sampled migrated record verdict harper-pro:5.1.230xD4×5 — no prefixRED harper-pro:5.2.0-beta.30xD4×5 — no prefixRED harper-pro:5.2.0-beta.40x42×5 — prefix presentGREEN, 11/11 gates So 5.1.23 is genuinely affected — the original conclusion holds and the scope claim (#1307 in every
v5.1.xfrom v5.1.2 onward) is unaffected, since it rested ongit merge-baserather than on that run. But the empirical row was not real when I posted it, and 5.1.14 and 5.1.23 are different assertions even when they happen to agree.beta.4 result, independently:
0x42on all five sampled records with all 11 gates green. That is external confirmation that #2014 fixes #2012 on an actual copied-filesmigrateOnStartmigration, on a published image, from outside the core repo.Also worth recording for anyone verifying this themselves: every API-level check passes on the RED builds.
COUNT(*)returns the full 3000/3000, a 5-id content spot check succeeds, a 6000-row cold read finds no null bodies, and__updatedtime__compares equal to the pre-migration v4 capture — because that value survives as a body attribute copied from v4. An earlier revision of my gate compared__updatedtime__viasearch_by_hashand passed on beta.3. Only reading the CF's raw bytes distinguishes the builds.
🤖 Posted by Claude Fable 5 on behalf of @heskew
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
On a database migrated LMDB→RocksDB via
storage.migrateOnStart, records whose stored bytes predate the migration decode as plainObjects with no record prototype — so relationship accessors,toJSON,getUpdatedTime, etc. are unreachable on exactly those records. Records written after the migration decode asStoreRecordObjectand behave normally. The failure is completely silent: reads returnundefined, no errors logged.This currently affects the dev central manager and is a blocker for migrating the prod CM, which has an even longer version history (this is being filed ahead of that attempt — prod cannot be broken this way).
Environment
harperfast/harper-pro:5.2.0-beta.3, RocksDB (migrated from LMDB viastorage.migrateOnStart)storage.migrateOnStartperformed the LMDB→RocksDB migration. The affected Organization record (2026-04-29) was written by an early-4.7-era build.loadAsInstance = falseresources, graphql schema with@relationshipsEvidence (live inspector evals on the dev CM process)
The table is healthy — struct prototype has all relationship getters, resolvers installed:
{"protoNames":["constructor","allowRead","roles","clusters","settings"], "attrRels":[{"name":"roles","hasResolve":true},{"name":"clusters","hasResolve":true},{"name":"settings","hasResolve":true}]}But per-record prototype linkage depends on when the record was last written:
StoreRecordObjectObject'clusters' in rec === false, readsundefinedObjectStoreRecordObjectDownstream, this produced:
GET /Organization/:idsilently missing itsclusters/settingsrelationship properties, androleData.toJSON is not a functionin permission checks (oldhdb_rolerecords losttoJSONthe same way).What does NOT reproduce it (important)
A clean-room reproduction of the same path does not show the bug: seed data on
harperdb/harperdb:4.7.33→ bootharper-pro:5.2.0-beta.3withstorage.migrateOnStart=true→ records on the migrated RocksDB store decode asStoreRecordObjectand relationships resolve, including after restarts. So the trigger involves something in the real instance's history. Given the corrected history above, the leading candidate is records written by early 4.7 betas / older 4.7.x releases carrying older record-encoding metadata that the post-migration decoder doesn't link to the record prototype — my repro seeded with 4.7.33 (late 4.7), which likely already wrote the newer encoding. Other candidates: the migration ran under an earlier 5.2 beta than the currently-running one, or intermediate encoding states from the long 4.7.x upgrade chain.Instance preserved for investigation
We are deliberately leaving the dev CM in the broken state (not running the rewrite workaround) so core can inspect affected records in situ — ping @Devin-Holland or devain for access/evals.
Observed workaround
Rewriting a record heals it (the constantly-patched Host records all decode correctly). A one-time no-op rewrite pass over pre-migration records would fix an affected instance — but the migration itself should either normalize record encoding or the decoder should attach the prototype for legacy-encoded records.
Asks
migrateOnStartre-encode (or the decoder tolerate) legacy records so long-history instances migrate safely.— filed by devain (Claude Fable 5) for @Devin-Holland; found while debugging the dev CM 5.2 upgrade (see central-manager#530 for the application-side fallout)