Repository navigation
search()/query() iterators run outside the request transaction: a handler cannot see its own writes, and its scans and point reads disagree about the snapshot #2506
Description
Activity
Reopening. This was closed as completed on 2026-09-11 with no linked PR, no referencing commit and no comment, and the mechanism is still present on
mainat68550c953. Two finding-triage passes (2026-09-14 and 2026-09-16) reached the same conclusion independently, and the second one found that the impact is worse than this issue currently describes.Still unfixed on today's main
resources/Table.ts:4628still opens the condition scan through a read transaction:const readTxn = txn.useReadTxn(target.snapshot === false);
DatabaseTransaction.useReadTxn→getReadTxnopens or reuses a nativeRocksTransaction, but staged writes live inthis.writes(an in-JS array) and are not applied to that native handle until commit — so even a reused read handle on the sameDatabaseTransactioncannot see the request's own pending writes. That is exactly what rocksdb-js#830 ("Store.getRange()silently ignoresoptions.transaction") describes, and #830 is still open.resources/Table.tshas changed by ~745 lines since55691f500(describe/estimate work, record-lock control entries, replication apply-failure handling) — none of it in the scan or delete path.resources/search.ts's only change is the #2187 options-object refactor, explicitly "no behaviour change". Of the two PRs cross-referenced here, #2556 changes scan lifetime under eviction, not which transaction a scan reads through, and #2498 is unrelated.Corroboration from this repo's own history: commit
4c353f880's body records thatunitTests/resources/rangeReadActivity.test.js's read-your-writes case "reproduces identically on origin/main6010ea9fd, harper#2506" — i.e. the defect was recorded as live in a commit message on the same day this issue was closed as done.The destroy half, which wasn't stated here before
The issue as written covers a handler not seeing its own writes. The condition-driven delete path turns that into data destruction, not just a missed row.
Table.ts:4145delete()takes the search-target branch:(scanTarget as any).select = ['$id']; for await (const entry of this.search(scanTarget)) { this._writeDelete((entry as any).$id); }
_writeDeletestages an unconditional delete-by-id. Its commit callback consults a prior same-transaction staged value only to clean up indices correctly (the narrower #1968 fix) — it never re-checks the record's current value against the delete's original predicate before removing it. So a row that an earlier staged write in the same handler moved out of the matching set is still on the stale candidate list and is deleted anyway.Both directions therefore hold:
direction effect under-delete a row staged into the matching set is never a candidate, so it survives a delete that should have removed it over-delete a row staged out of the matching set is deleted regardless, because nothing re-checks the predicate Engine coverage
The compound-AND and primary-key extension checks out structurally:
transformToEntries(Table.ts:7366) andloadLocalRecord(:7029) reload viaprimaryStore.getEntry(id, {transaction}), a point lookup — a materially different path fromgetRange, which is the asymmetry rocksdb-js#830's own title draws. So the AND-sibling protection that exists is RocksDB-only.On LMDB there is none at all:
resources/LMDBTransaction.ts:41getReadTxnreturns a plain lmdb-js read-only MVCC transaction with no visibility into pending writes for any access pattern, point or range. That half is lower urgency given LMDB is deprecated — but RocksDB alone still carries both directions above, which is what makes this worth reopening rather than downgrading.Test anchor that already exists
unitTests/resources/rangeReadActivity.test.js's "preserves read-your-writes through primary and index searches" (added 2026-09-09,bf0d17d78) is an in-tree anchor for this and is reportedly red on plainmain. Whoever scopes the fix should expect to turn it green. Two gaps worth knowing about: it is gatedif (isLMDB) this.skip(), and it exercises onlysearch()visibility — not the condition-driven delete/update write-path consequence above, which nothing in the suite currently covers.If this was closed deliberately — a fix landed somewhere I couldn't find, or the behaviour is intended — linking that here would settle it; I could not find it from the timeline, the commit log, or the code.
- added a commit that references this issue
on Sep 16, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
Every iterator read surface of a request transaction runs outside that transaction. A handler
that writes a row and then reads it back through
search()/query()— index scan, full scan, oreven a condition on the primary key — does not see its own write. Only the point read
(
tables.X.get(id)) does.The same gap has a second consequence: those scans are not on the transaction's snapshot either, so
a row a different request commits while the transaction is open is visible to its full scan and
is not visible to its point read. One request's scans and point reads disagree about the state
of the database.
Reported from a live 5.2.7 deployment as a secondary-index bug. The secondary index is not what is
special — it was simply the surface the application used.
Measured
One request transaction, one staged write, four read surfaces (RocksDB, reproduced on
main):Control, in the same transaction and the same index search: three rows committed before the
transaction opened are returned. So the index read works; only the transaction's own write is
missing from it.
Snapshot consequence — a row committed by another request 750 ms into an open transaction:
Waiting inside the transaction changes nothing (the reporter held it 10 s); a staged write's
visibility is not time-dependent.
LMDB misses the staged write on all four surfaces, point read included, which matches the
documented LMDB staging behaviour (
DESIGN.md:117, HarperFast/documentation#531). The get-vs-scansplit is RocksDB-only.
Root cause
Below the Table layer, with the same read handle
search()uses(
Table._readTxnForContext(context)):rocksdb-js
Store.getRange()builds the native iterator fromthis._contextand never readsoptions.transaction, whileget/getSyncroute it throughgetTxnId(). Filed as the root causein HarperFast/rocksdb-js#830.
Harper's query engine passes the option on every range read that should be transactional:
resources/search.ts—executeConditions()resolvestable._readTxnForContext(context)andhands it to
searchByIndex(), which passes it asindex.getRange({ …, transaction }); theprimary-key and full-scan paths do the same.
resources/Table.ts—_readTxnForContext()returnstxnForContext(context).getReadTxn(),i.e. the transaction the handler's writes are staged into. The write really is in the native
transaction; the point read proves it. Only the iterator cannot see it.
Impact
Silent.
search()returns fewer rows, with no error and nothing in the log, so it isindistinguishable from "those rows do not exist". The reporting application enforces a per-account
row cap by writing and then counting through an indexed search, and overshoots under concurrency —
6 concurrent adds against a limit of 3 leave 4 rows, 3 trials of 3. Any tenancy, quota,
dedupe, or state-machine check that writes and then scans in one request has the same exposure, and
so does any handler that assumes its scans and its point reads describe the same snapshot.
Not a data-corruption bug: committed state is correct, and a fresh request reads it correctly.
Reproduction
A fixture reproducing this is written but not pushed, so treat the recipe below as the spec rather
than looking for the branch. One custom
Resourceendpoint, inside its single request transaction:await tables.Item.put({ id, groupId, … })on a table whosegroupIdis@indexed;tables.Item.get(id),search({conditions:[{attribute:'id', value:id}]}),search({conditions:[]})(full scan),search({conditions:[{attribute:'groupId', value:groupId}]});groupIdfrom earlier requests first, and assert they ARE returned bystep 2's index search — the control that the index read works inside this transaction.
For the snapshot half: give the endpoint a sleep between the write and the reads,
POSTa secondrequest that commits a row under the same
groupIdwhile it sleeps, and assert the scans and thepoint read agree about whether that row is visible.
Fix shape
Two spellings of the same change:
Store.getRange()in rocksdb-js sooptions.transactionmeans what it means onget(); ortransaction.getRange()in Harper — the transaction as the iterator's context — whichalready works today.
Either way every transactional scan starts pinning its read transaction for the life of the
iterator, which lands squarely in
DatabaseTransaction's read-transaction lifetime accounting(
readTxnsUsed,doneReadTxn, and the "read iterators held a committed transaction's snapshot pastthe open-transaction limit" path). Scoping that is the substance of this issue; the rocksdb-js
option being inert is the substance of the other.
Not this
delete K; put Kstripping a live record's index entries) is adifferent defect. Its fix, Fix a replicated delete+put of the same record dropping all of its secondary index entries #2235, first shipped in v5.2.4 and is present in the 5.2.7 that
reproduces this — verified by
merge-base --is-ancestoragainst that release'scorepointer.finding that concurrent
create()on one primary key admits a nondeterministic number of racersis Add ability to exclusively lock a record for performing a safe operation #483's territory.
tables.X.get(id)read-your-writes semantics per engineand says nothing about
search(); it should pick up whatever this resolves to.Versions
Reproduced on
mainand observed live on harper-pro 5.2.7 (coref883a3bf). RocksDB only for theget-vs-scan split; the no-own-write-on-scans half applies to both engines.