Skip to content

search()/query() iterators run outside the request transaction: a handler cannot see its own writes, and its scans and point reads disagree about the snapshot #2506

Description

@kriszyp

Summary

Every iterator read surface of a request transaction runs outside that transaction. A handler
that writes a row and then reads it back through search()/query() — index scan, full scan, or
even a condition on the primary key — does not see its own write. Only the point read
(tables.X.get(id)) does.

The same gap has a second consequence: those scans are not on the transaction's snapshot either, so
a row a different request commits while the transaction is open is visible to its full scan and
is not visible to its point read. One request's scans and point reads disagree about the state
of the database.

Reported from a live 5.2.7 deployment as a secondary-index bug. The secondary index is not what is
special — it was simply the surface the application used.

Measured

One request transaction, one staged write, four read surfaces (RocksDB, reproduced on main):

ownWrite: { pointGet: true, byPrimaryCondition: false, fullScan: false, byIndex: false }

Control, in the same transaction and the same index search: three rows committed before the
transaction opened are returned. So the index read works; only the transaction's own write is
missing from it.

Snapshot consequence — a row committed by another request 750 ms into an open transaction:

{ pointGet: false, fullScan: true, byIndex: false }

Waiting inside the transaction changes nothing (the reporter held it 10 s); a staged write's
visibility is not time-dependent.

LMDB misses the staged write on all four surfaces, point read included, which matches the
documented LMDB staging behaviour (DESIGN.md:117, HarperFast/documentation#531). The get-vs-scan
split is RocksDB-only.

Root cause

Below the Table layer, with the same read handle search() uses
(Table._readTxnForContext(context)):

store.get(id,      { transaction })          → true    (write batch consulted)
store.getRange({ start, end, transaction })  → false   (option accepted and dropped)

rocksdb-js Store.getRange() builds the native iterator from this._context and never reads
options.transaction, while get/getSync route it through getTxnId(). Filed as the root cause
in HarperFast/rocksdb-js#830.

Harper's query engine passes the option on every range read that should be transactional:

  • resources/search.ts — executeConditions() resolves table._readTxnForContext(context) and
    hands it to searchByIndex(), which passes it as index.getRange({ …, transaction }); the
    primary-key and full-scan paths do the same.
  • resources/Table.ts — _readTxnForContext() returns txnForContext(context).getReadTxn(),
    i.e. the transaction the handler's writes are staged into. The write really is in the native
    transaction; the point read proves it. Only the iterator cannot see it.

Impact

Silent. search() returns fewer rows, with no error and nothing in the log, so it is
indistinguishable from "those rows do not exist". The reporting application enforces a per-account
row cap by writing and then counting through an indexed search, and overshoots under concurrency —
6 concurrent adds against a limit of 3 leave 4 rows, 3 trials of 3. Any tenancy, quota,
dedupe, or state-machine check that writes and then scans in one request has the same exposure, and
so does any handler that assumes its scans and its point reads describe the same snapshot.

Not a data-corruption bug: committed state is correct, and a fresh request reads it correctly.

Reproduction

A fixture reproducing this is written but not pushed, so treat the recipe below as the spec rather
than looking for the branch. One custom Resource endpoint, inside its single request transaction:

  1. await tables.Item.put({ id, groupId, … }) on a table whose groupId is @indexed;
  2. read the row back four ways and report each as a boolean —
    tables.Item.get(id),
    search({conditions:[{attribute:'id', value:id}]}),
    search({conditions:[]}) (full scan),
    search({conditions:[{attribute:'groupId', value:groupId}]});
  3. seed rows under the same groupId from earlier requests first, and assert they ARE returned by
    step 2's index search — the control that the index read works inside this transaction.

For the snapshot half: give the endpoint a sleep between the write and the reads, POST a second
request that commits a row under the same groupId while it sleeps, and assert the scans and the
point read agree about whether that row is visible.

npm run build
npm run test:integration -- "integrationTests/database/index-read-your-writes.test.ts"
HARPER_STORAGE_ENGINE=lmdb npm run test:integration -- "integrationTests/database/index-read-your-writes.test.ts"

Fix shape

Two spellings of the same change:

  1. fix Store.getRange() in rocksdb-js so options.transaction means what it means on get(); or
  2. iterate transaction.getRange() in Harper — the transaction as the iterator's context — which
    already works today.

Either way every transactional scan starts pinning its read transaction for the life of the
iterator, which lands squarely in DatabaseTransaction's read-transaction lifetime accounting
(readTxnsUsed, doneReadTxn, and the "read iterators held a committed transaction's snapshot past
the open-transaction limit" path). Scoping that is the substance of this issue; the rocksdb-js
option being inert is the substance of the other.

Not this

Versions

Reproduced on main and observed live on harper-pro 5.2.7 (core f883a3bf). RocksDB only for the
get-vs-scan split; the no-own-write-on-scans half applies to both engines.

Activity

  1. added this to the v5.3 milestone on Sep 4, 2026
  2. added theissue type on Sep 4, 2026
  3. self-assigned this
    on Sep 4, 2026
  4. kriszyp commented on Sep 16, 2026

    @kriszyp
    MemberAuthor

    Reopening. This was closed as completed on 2026-09-11 with no linked PR, no referencing commit and no comment, and the mechanism is still present on main at 68550c953. Two finding-triage passes (2026-09-14 and 2026-09-16) reached the same conclusion independently, and the second one found that the impact is worse than this issue currently describes.

    Still unfixed on today's main

    resources/Table.ts:4628 still opens the condition scan through a read transaction:

    const readTxn = txn.useReadTxn(target.snapshot === false);

    DatabaseTransaction.useReadTxn → getReadTxn opens or reuses a native RocksTransaction, but staged writes live in this.writes (an in-JS array) and are not applied to that native handle until commit — so even a reused read handle on the same DatabaseTransaction cannot see the request's own pending writes. That is exactly what rocksdb-js#830 ("Store.getRange() silently ignores options.transaction") describes, and #830 is still open.

    resources/Table.ts has changed by ~745 lines since 55691f500 (describe/estimate work, record-lock control entries, replication apply-failure handling) — none of it in the scan or delete path. resources/search.ts's only change is the #2187 options-object refactor, explicitly "no behaviour change". Of the two PRs cross-referenced here, #2556 changes scan lifetime under eviction, not which transaction a scan reads through, and #2498 is unrelated.

    Corroboration from this repo's own history: commit 4c353f880's body records that unitTests/resources/rangeReadActivity.test.js's read-your-writes case "reproduces identically on origin/main 6010ea9fd, harper#2506" — i.e. the defect was recorded as live in a commit message on the same day this issue was closed as done.

    The destroy half, which wasn't stated here before

    The issue as written covers a handler not seeing its own writes. The condition-driven delete path turns that into data destruction, not just a missed row.

    Table.ts:4145 delete() takes the search-target branch:

    (scanTarget as any).select = ['$id'];
    for await (const entry of this.search(scanTarget)) {
        this._writeDelete((entry as any).$id);
    }

    _writeDelete stages an unconditional delete-by-id. Its commit callback consults a prior same-transaction staged value only to clean up indices correctly (the narrower #1968 fix) — it never re-checks the record's current value against the delete's original predicate before removing it. So a row that an earlier staged write in the same handler moved out of the matching set is still on the stale candidate list and is deleted anyway.

    Both directions therefore hold:

    direction effect
    under-delete a row staged into the matching set is never a candidate, so it survives a delete that should have removed it
    over-delete a row staged out of the matching set is deleted regardless, because nothing re-checks the predicate

    Engine coverage

    The compound-AND and primary-key extension checks out structurally: transformToEntries (Table.ts:7366) and loadLocalRecord (:7029) reload via primaryStore.getEntry(id, {transaction}), a point lookup — a materially different path from getRange, which is the asymmetry rocksdb-js#830's own title draws. So the AND-sibling protection that exists is RocksDB-only.

    On LMDB there is none at all: resources/LMDBTransaction.ts:41 getReadTxn returns a plain lmdb-js read-only MVCC transaction with no visibility into pending writes for any access pattern, point or range. That half is lower urgency given LMDB is deprecated — but RocksDB alone still carries both directions above, which is what makes this worth reopening rather than downgrading.

    Test anchor that already exists

    unitTests/resources/rangeReadActivity.test.js's "preserves read-your-writes through primary and index searches" (added 2026-09-09, bf0d17d78) is an in-tree anchor for this and is reportedly red on plain main. Whoever scopes the fix should expect to turn it green. Two gaps worth knowing about: it is gated if (isLMDB) this.skip(), and it exercises only search() visibility — not the condition-driven delete/update write-path consequence above, which nothing in the suite currently covers.

    If this was closed deliberately — a fix landed somewhere I couldn't find, or the behaviour is intended — linking that here would settle it; I could not find it from the timeline, the commit log, or the code.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Fields

Priority

P1

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions