Repository navigation
Mid-log corrupt transaction-log frame silently truncates replay and replication — acknowledged writes lost #2016
Description
Activity
- addedbugSomething isn't workingSomething isn't workingarea:storageStorage engine, LMDB/RocksDB, compactionStorage engine, LMDB/RocksDB, compactionarea:replicationReplication, clusteringReplication, clustering
on Jul 31, 2026 - added a parent issue
on Aug 6, 2026 - added a commit that references this issue
on Aug 6, 2026 Status on this, since the fix is split across three repos and only part of it can land today.
The reader half — rocksdb-js#750 — fix(txnlog): resync past a mid-log corrupt frame instead of ending the log. Ready, green on every platform/runtime leg, awaiting human review.
query()now reports a break as aCorruptFrameErrorcarryingresyncPosition(where valid framing resumes) plus the unreadable byte count, and advances its reader there before throwing — so only the torn frame is lost, not everything behind it. This is suggested direction 2. A 12-entry log with a broken second frame yields 11 entries where the old reader yielded 1.The consumer half — harper#2087 — fix(replay): resync past a mid-log corrupt transaction-log frame and surface it as data loss. Draft, rebased onto main today.
endIteratorOnCorruptFramekeeps pulling when the error carries a resume point instead of latching, capped at 32 resyncs per iteration; a mid-log break now logs aterrorand says entries were lost, which is suggested direction 1's severity half. Kept a draft deliberately: no released rocksdb-js setsresyncPosition(verified against 2.7.0), so it is inert until a bump ships — behavior is unchanged today, which is why the two need not land in lockstep, but also why neither closes this on its own.The health signal — harper-pro#667 — cluster_status reports a stream healthy after it has lost transaction-log entries. Filed just now, not started. harper#2087 accumulates break sites and exposes them as
getCorruptFrameReports()(location, mid-log vs torn tail, unreadable bytes, whether iteration stopped, occurrence counts), but nothing consumes it yet, andcluster_statusis harper-pro. Until that is wired, the only operator-visible signal is still a log line — the exact gap that let this run 2.2 days here and 11 days in #2063. That covers suggested directions 1 (surfacing) and 4.Suggested direction 3 — don't acknowledge a write whose append failed — is not addressed and is the only one that prevents rather than recovers. Filed as rocksdb-js#748 — A failed transaction-log append orphans its partial bytes, baking a mid-file framing break into the log. Related: rocksdb-js#749 — Uncommitted transaction-log reads bound corruption checks by mapped capacity, which is why a torn frame with a plausible declared length still goes undetected on the boot-replay path.
One honest coverage gap on both PRs: every test drives synthetic buffers or iterators. Nothing reopens a genuinely damaged log on disk and watches a real consumer replicate past the break — that test lives in harper-pro and is not written.
🤖 Claude Opus 5
The end-to-end test is now up as harper-pro#670 — test(cluster): prove mid-log txnlog tear recovery end-to-end, which closes the coverage gap noted above: a real two-node stream reading a genuinely damaged log off disk, rather than synthetic buffers.
It confirms the fix (39/60 rows on released rocksdb-js 2.7.0, 60/60 on the #750 build) and turned up a residual: the readable tear shape — the one a partial append usually leaves — still wedges the receiver even on a fixed engine, because the torn frame is yielded as a well-formed entry with a garbage payload and no reader can tell. Filed as harper-pro#669; details in my note on #2087.
So suggested direction 2 (frame resync) is proven for the unreadable shape, and direction 3 (don't acknowledge a write whose append failed — rocksdb-js#748) is now the load-bearing one, since prevention is what removes the poison entry rather than surviving it.
🤖 Claude Opus 5
- added a commit that references this issue
on Aug 19, 2026 A data point for this issue's "corrupt frame silently accepted" premise, from dispatch QA finding F-289: rocksdb-js's open-time recovery scan has no payload checksum.
scanTransactionLogForRecovery()(src/binding/transaction_log/transaction_log_recovery.h, rocksdb-js origin/main a941a670) classifies only framing integrity —Clean/TruncateTail/MidFileCorruption— and the header docs state outright that a payload-level bit-flip is "indistinguishable... without a checksum". So a corruption that leaves the frame header and length prefix intact but flips payload bytes classifies asCleanand is loaded/replayed as if valid — exactly the silent-truncation-vs-silent-bad-data hazard this issue is about, one layer deeper thanreplayLogsGuards.ts'sendIteratorOnCorruptFrame(which only fires on framing breaks).This came up while investigating a separate, still-unconfirmed claim (a 2-byte payload flip appearing to wedge restart); that wedge half is being localized separately and may be a harness signal-handling artifact. But the payload-checksum gap itself is real and independent of the wedge question, and it argues for a frame-level payload checksum (or CRC) as part of hardening this path.
From dispatch QA finding F-289. — Claude (Fable 5)
- added 7 commits that reference this issue
on Aug 31, 2026 Root causes have been addressed, these are pretty rare edge cases at this point.
Data point for the "cases that look like this issue": a production occurrence on 5.2.6 (rocksdb-js 2.7.1) with the exact
declared length N overruns the log (limit=L)fail-stop, where there was no corrupt frame at all.Lequalled the pinned committed watermark on each of three affected nodes (full table in harper#2073, comment of 2026-10-01). Replication to every peer stopped at the first entry straddling that byte,cluster_statusstayedconnected: trueand replication metrics stayed at 0 lag for 8 days, matching the "indefinite replication starvation" item here.Worth splitting the two causes in the reader's handling: a frame that overruns the watermark bound is not evidence of corruption and should not fail-stop the stream, whereas a frame that overruns the file is.
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
A mid-log corrupt frame in a table's local transaction log silently truncates every reader of that log — crash-recovery replay and replication — at the position of the tear. Because tables run WAL-less by default, replay is the only thing standing between an unclean exit and the loss of acknowledged writes, so a single torn frame converts into:
cluster_statusand component health stay green.warnlevel, latched to fire once per log.The design assumption that breaks
endIteratorOnCorruptFrame(resources/replayLogsGuards.ts) deliberately treats a corrupt frame as end-of-log:That is sound when the tear is at the tail — a torn final write, nothing after it. It is wrong when the tear is mid-log: a failed append (ENOSPC/EDQUOT partial write) that the process survives. The process keeps accepting writes and appending valid frames after the tear, every one of them acknowledged to clients — and every one of them unreachable to any future reader.
With
options.disableWAL ??= true(resources/databases.ts:163), those post-tear writes have exactly two homes: the memtable and the txnlog-beyond-the-tear. An unclean exit destroys the first and the tear amputates the second.Observed incident (5.1.x, two-node hosted cluster)
position 3bc071 of log 23) and repeated at every subsequent boot.Connected,cluster_statussaidreplicates: true) but moved zero rows — each new reader re-stopped at the same frame. Nothing distinguished this from healthy replication except manually diffing record counts across nodes.The records were recoverable only because the operator happened to hold an off-box copy.
Suggested directions
Ordered roughly by how much of the problem each removes:
error, surface it in health/cluster_status, and don't let replay conclude "done".Related
RangeErroron the same class of torn entry (same incident family; that one crashes the caller, this one silently truncates).🤖 Filed by Claude on behalf of @heskew