Skip to content

Validate block-store responses before retry - #1411

Merged
DZakh merged 5 commits into
claude/block-store-reorg-tracking-2ucqi4from
codex/block-store-response-retry
Jul 14, 2026
Merged

DZakh merged 5 commits into
claude/block-store-reorg-tracking-2ucqi4from
codex/block-store-response-retry

Conversation

@DZakh

@DZakh DZakh commented Jul 13, 2026

Copy link
Copy Markdown
Member

Summary

  • validate page-internal and cross-response block hash consistency in the Rust block store
  • retry inconsistent block responses through the existing onReorg path while preserving request statistics
  • move EVM and SVM rollback hash pagination into Rust, including parent-slot and parent-blockhash validation for skipped SVM slots
  • materialize block data from persistent block stores across source implementations

Why

A response can contain block hashes from inconsistent backend instances or forks. Keeping response conflict metadata and cross-store comparison inside BlockStore distinguishes an invalid response from a confirmed chain reorg and keeps block-only data out of ReScript.

Validation

  • cargo check -p envio --lib
  • cargo fmt --all -- --check
  • pnpm build in packages/envio
  • pnpm build in scenarios/test_codegen
  • focused Vitest suite: 67 tests passed
  • Rust unit suites: BlockStore 21, EVM HyperSync 8, SVM query 4, request stats 1
  • git diff --check

@coderabbitai

coderabbitai Bot commented Jul 13, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 6af1b9d1-abc0-46d2-ae16-416fe19e49f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

Comment on lines +135 to +139
if to_block - from_block > 1_000 {
return Err(map_err(anyhow::anyhow!(
"Invalid block data request. Range of block numbers is too large. Max range is 1000. Requested range: {from_block}-{to_block}"
)));
}

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't see reason for that imitation.

let started = Instant::now();
let response = self.get_raw(query).await;
request_stats.push(RequestStat {
method: "getBlockHashes".to_string(),

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

queryBlockHashes

// A successor can only prove a missing suffix. Interior omissions
// already have a returned descendant, so another later block cannot
// repair their parent link.
if aggregate.missing_hashes(block_numbers).contains(&to_slot) {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Explain why this is needed

…:enviodev/hyperindex into codex/block-store-response-retry

# Conflicts:
#	packages/cli/src/evm_hypersync_source/mod.rs
#	packages/cli/src/evm_rpc_source/mod.rs
#	packages/envio/src/sources/HyperSyncClient.res
@DZakh
DZakh marked this pull request as ready for review July 14, 2026 10:51
// the upper slot was skipped or leaves it missing as an invalid
// response; without this lookup a valid skipped upper slot would be
// retried forever.
if upper_slot_needs_successor_proof {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If HyperSync didn't return a slot in the range between from slot and response slot then it means it is trully missing.

@DZakh
DZakh merged commit ffa26ac into claude/block-store-reorg-tracking-2ucqi4 Jul 14, 2026
8 checks passed
@DZakh
DZakh deleted the codex/block-store-response-retry branch July 14, 2026 14:04
DZakh added a commit that referenced this pull request Aug 13, 2026
* Move reorg tracking from the ReScript registry into the Rust block store

The per-chain BlockStore now owns reorg detection: merging a fetch-response
page compares block hashes and reports the lowest in-threshold mismatch
(discarding the page in rollback mode, overwriting in detect-only mode),
pruning keeps in-threshold hashes as hash-only rows, and rollback reads
(getHash, getHashedBlockNumbers, latestValidBlock) replace the JS-side
ReorgDetection registry.

- Fuel gets a first-class store (height/time/id) built by the Rust client;
  Fuel blocks are materialised from the store instead of carried inline.
- BlockStore.fromJs builds pages from sparse JS blocks: RPC contributes
  hash-only observations, simulate an empty page, and stored reorg
  checkpoints seed the store on resume.
- The EVM HyperSync rollback-guard blocks are inserted into the page store
  on the Rust side; sources no longer return a separate blockHashes array.
- ReorgDetection.res shrinks to the shared data types and log params.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Preserve registry reorg semantics in the store and port the test suite

- Store detection hashes as their JS string form for every ecosystem, so
  pages built from JS observations compare byte-for-byte with fetched blocks
  (and mismatch reports return the original strings).
- Record within-page hash conflicts (the same block observed twice with
  different hashes in one response) while a page is built and report them
  from merge, matching the old duplicate-collision detection.
- Rollback keeps hash-only rows on non-reorg chains — their scanned hashes
  stay valid while refetch repopulates the data — and drops everything above
  the target on the reorg chain.
- Checkpoint block hashes are gated by the chain's own reorg threshold
  (sourceBlockNumber - maxReorgDepth) instead of the global flag.
- Port ReorgDetection/SourceBlockHashes/ChainState/rollback tests to the
  store-backed API and pin Fuel.blockFields against the Rust ordering.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Address review: Fuel id field first, fold page conflict into store lock

- FuelBlockField orders id first (Fuel.res blockFields matches).
- The within-page hash conflict lives inside the store's single Mutex
  alongside the table instead of a second lock; it stays on the struct
  because a page is built across several insert calls (response blocks,
  then guard rows) and merge reads it later.
- Clarify why the EVM hash column is filled outside evm_block_col.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Store detection hashes as validated hex bytes; Fuel order height,id,time

EVM/Fuel block hashes are hex-validated and stored as bytes again: fromJs
pages reject non-hex hashes (e.g. arbitrary marker strings) with a
validation error instead of storing them opaquely. The hash column is
variable-width — 32 bytes for fetched blocks — so hex test fixtures can
stay short. Test mocks now use valid even-length hex hashes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Validate block-store responses before retry (#1411)

* Validate block-store responses before retry

* Address block hash query review feedback

* Track SVM cursor coverage in block stores

* Fix SVM parent validation and metrics

* Address block-store reorg review findings: config authority, test resilience, seam validation (#1442)

* Address review findings: threshold authority, EVM parent-link check, test coverage

- Use the resumed-from-DB maxReorgDepth and per-chain shouldRollbackOnReorg
  everywhere (Batch checkpoints, getHighestBlockBelowThreshold, reorg
  logging/rollback decision) instead of mixing them with config values
- Validate EVM parent links in response stores: block N's parentHash must
  match block N-1's hash, within a page and across page seams; select
  parentHash in the EVM getBlockHashes re-fetch (Fuel has no parent-id
  field, so the check stays EVM/SVM-only)
- Make shouldRollbackOnReorg/maxReorgDepth required in ChainState.make
- Cover ChainState threshold arithmetic (depth changes across restarts,
  clamping, registerReorgGuard boundary) and applyBatchProgress hash
  retention with new tests
- Harden field_table: hard width checks on fixed columns, 64-field cap
- Drop stale ReorgDetection comments, dedupe native-failure unpacking,
  map rate-limited errors in the SVM getBlockHashes path too, fix test
  indentation

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Surface EVM block-hash no-progress as an error for SourceManager to retry

Matches the SVM path: instead of silently sleeping 100ms in-process
(unbounded, no logging, no failover), the paginator returns a
RequestFailed error carrying the accumulated request stats. SourceManager
logs each retry with backoff and switches to another source on repeated
failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Friendlier no-progress message and a fast first retry for block-hash fetches

The replica-drift error now explains itself (routing to a replica slightly
behind the head, safe to continue after a retry), and SourceManager's
first block-hash retry backs off only 100ms to match how quickly a lagging
replica usually catches up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Polish block-hash retry: doubling backoff, message-first logging

- Backoff doubles from 100ms (capped at 60s) instead of stepping by 1s
- The native failure's own message becomes the retry log's msg (the
  replica-drift text is self-explanatory); generic failures keep the
  err payload
- Native failure causes are plain JS errors now, so logs no longer show
  a NativeRequestFailed wrapper
- Drop the logType field from the block-hash query logger and reword the
  replica-drift message

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Shorten the replica-drift message

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Use the resumed reorg depth for the pre-threshold fetch lag

Codex review: with a reduced configured depth, blockLag from the config
value would let fetching enter the stored rollback window without
history while detection still compares the resumed window.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Fail test runs with the source error instead of retrying forever

In production a failing source is retried indefinitely (backoff and
failover keep the indexer alive), which in tests turns an unreachable
endpoint into a bare test-runner timeout with no context. SourceManager
now accepts a maxRetries cap (ENVIO_MAX_SOURCE_RETRIES) enforced across
the height, getItems, and getBlockHashes retry loops; when exhausted the
run fails with the underlying error. The test indexer worker defaults
the cap to 1, and generated templates give vitest 60s so the real error
surfaces before the runner's axe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Make live-endpoint tests resilient to hung connections

The test indexer worker now caps the HyperSync request timeout at 10s
(production default is 120s, which outlives every test timeout and turns
a hung connection into a context-free failure) so a hang fails fast and
retries on a fresh connection. The scenarios suite retries a failed test
once on CI - tests run sequentially with a fresh indexer per test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Reject EVM block-hash pages that don't cover their range densely

Codex review: the parent-link check skips absent neighbours (event pages
are sparse by design), so an include_all_blocks page omitting an interior
block or a parentHash could hide a mixed-fork seam. Each page is now
validated to carry every block in its covered range with hash and parent
hash before it joins the aggregate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Honor the retry cap in the subscription fallback poller

Codex review: the stale-subscription fallback swallowed every
getHeightOrThrow failure, so with a subscription installed the cap never
fired and a dead endpoint could still hang a test run. The fallback's
rejection is separately observed since it can lose the height race.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Retry a template smoke test once before failing the job

The template suites index real blocks through live HyperSync; one hung
connection on the runner shouldn't fail the job now that a failed
attempt surfaces quickly with a real error.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Retry e2e smoke tests once on CI

Same rationale as the scenarios suite: the smoke tests hit live
HyperSync, and a hung runner connection now fails fast with a real
error, so a single retry absorbs it without hiding deterministic
failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Replace the global CI retry with per-test retries on live-endpoint tests

Only the tests that hit live HyperSync/RPC endpoints retry (following the
existing SourceBlockHashes pattern of {retry: 3}): the HyperSync client
live tests, the RPC height check, the createTestIndexer tests that fetch
mainnet, and the e2e smoke test. Deterministic tests fail on the first
attempt again. The Vitest binding options gained a timeout field so the
corrupted-token test keeps its 60s budget alongside the retry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Keep the test-worker HyperSync timeout at 30s

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Widen template harness timeout padding from 10s to 30s

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Simplify test-run resilience knobs

Drop the template harness re-run hack and the derived outer timeout (the
config value is the single budget now), run template suites with a 30s
per-test vitest timeout, and rely on ENVIO_MAX_SOURCE_RETRIES=3 in test
workers instead of overriding the HyperSync client timeout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

* Replace parent-hash validation with hash-collision checks

A single HyperSync response is internally consistent, so parent-link
validation only guarded the seams between paginated block-hash requests.
Those seams are now covered directly: each follow-up page re-requests the
last returned block, and a fork switch between requests surfaces as a
hash collision on the overlapping block via the existing duplicate
detection. The EVM/SVM parent-link checks, the dense-range validation,
and the parent fields in block-hash queries are gone.

RPC responses can mix forks since every block is fetched separately, so
getBlockHashes now derives a minimal (number-1, parentHash) row from each
block - like the items path already did - letting the page's existing
collision check cross-validate separately fetched blocks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019RcUijSCoEY19MCEbdh6fL

---------

Co-authored-by: Claude <noreply@anthropic.com>

* Fix ChainState.make call site after merging required reorg params

The review-findings commit made ~shouldRollbackOnReorg/~maxReorgDepth
required; the CrossChainState_test call site introduced by the query-sizing
merge needed ~shouldRollbackOnReorg added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Address block-store reorg review: fixed-32 hashes, immutable batch snapshot, cleanups (#1448)

- field_table: extract shared row-reduction/drop helpers from
  prune_keeping_field and rollback_keeping_field.
- block_store: trim EvmBlockInput to the (number, timestamp, hash) trio the
  JS callers actually send, dropping the unreachable full-block conversion.
- block_store: store the EVM hash as a fixed 32-byte column again; fromJsEvm
  left-pads shorter mock markers into that width (a no-op for real 32-byte
  observations), so storage stays fixed-width without input-fixture churn.
- block_store: extract a ResponsePage type owning the response-only state
  (within-response conflict + SVM cursor coverage), so the persistent store no
  longer carries it and merge resets it explicitly instead of clearing fields.
- request_stats/Source: mark the native-failure envelope with an explicit
  ENVIO_NATIVE_FAILURE: prefix so ReScript decodes only our own payloads and
  never a coincidental JSON error message.
- Batch/ChainState: snapshot the in-threshold scanned hashes when the batch is
  assembled instead of reading the live block store per block; drops the live
  blockStore/maxReorgDepth off chainBeforeBatch and removes the per-block napi
  getHash calls from the checkpoint loop.
- tests: migrate the getHash-derived EVM hash expectations to the padded
  32-byte form via a shared MockIndexer.evmBlockHash helper.


Claude-Session: https://claude.ai/code/session_01K3yMGs2o2DbQqu6LSTcHEd

Co-authored-by: Claude <noreply@anthropic.com>

* Address block-store reorg review: drop rollback data fully, trim guard/SVM plumbing (#1450)

- Rollback now always drops all data including hashes; the refetched range
  repopulates them, so stale hashes can't linger for reorg detection. Removes
  the keepHashes flag and the now-dead rollback_keeping_field.
- Stop returning the rollback guard from the EVM event response; its blocks are
  already inserted into the BlockStore on the Rust side. The head timestamp now
  comes from the last item, matching the Fuel source path.
- Drop the unused height/parent_slot/parent_hash fields from SvmBlockInput; JS
  only observes slot/time/hash for reorg tracking.
- Reword the inconsistent-response retry log to describe a partial reorg
  indicator instead of an "internally inconsistent response".


Claude-Session: https://claude.ai/code/session_012JJNrNK67GrjneNK8QFAhh

Co-authored-by: Claude <noreply@anthropic.com>

* Centralize source retry policy, drop the retry-cap env var

Reorg-detection state is rebuilt from scratch after a rollback, so the
comment claiming BlockStore.rollback gates hash retention on isReorgChain
described a parameter that never existed. Removed.

Move the "backend instance hasn't reached this block yet" retry out of the
per-ecosystem sources and into SourceManager. EVM and Fuel each built their
own backoff schedule and message for it, and the native block-hash paginator
reported it as free-text that SourceManager could only treat as an unknown
failure. They now raise Source.SourceBehindHead — the Rust clients through a
SOURCE_BEHIND_HEAD:<block> marker, alongside the existing RATE_LIMITED: one —
so one policy covers getItems and getBlockHashes across EVM, SVM and Fuel.

Extract the failover decision that only the WithBackoff path implemented into
backoffBeforeRetry, and reuse it for the behind-head and inconsistent-response
retries. Marking lastFailedAt only demotes a source in the selection order, so
with no alternative to move to the previous inconsistent-response path skipped
its delay and spun; those retries now carry a 50ms floor. Callers that ask for
no backoff keep it.

Drop ENVIO_MAX_SOURCE_RETRIES and the maxRetries plumbing: nothing set the
variable, and its six call sites turned a routine head-of-chain condition into
a run-ending error.

Keep the rollback guard's block timestamp in the block store instead of
storing the head block hash-only, so the row materializes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RrYmidxrL75hDhP24nXmd6

* Give SVM a real reorg depth, delete the write-only fetched-block timestamp

SVM ran with maxReorgDepth 0, so the comparison window was [knownHeight, ∞)
and the block-hash machinery could only ever fire on the head slot. Tower BFT
roots a block once 32 votes lock it in, which is the protocol's own bound on
how far a fork can be replaced — but that bound counts blocks while the
threshold is measured in slot numbers, and skipped slots make the slot
distance the larger of the two. Default to 64 so the window still spans 32
blocks at any realistic skip rate, and honour a configured value instead of
discarding it. Fuel keeps 0; nothing about its finality is established here.

latestFetchedBlockTimestamp was never read. Five sources computed it,
ChainFetching packed it into FetchState, and FetchState itself wrote a literal
0 into the same field in eight places. Progress latency reads the materialized
event block's timestamp instead, ordering is block-number based, and no
column or API exposes it. Removing it also retires the per-page block-header
lookups that existed only to feed it, and the record it lived on loses its
second field — hence blockRef.

Block timestamps that reach handlers are untouched: they come from the block
store, which now also keeps the rollback guard's head timestamp.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RrYmidxrL75hDhP24nXmd6

* Cover reorg rollback against a real JSON-RPC server

The rollback suite drives sources that fabricate their own responses, so
nothing exercised the hashes a real source harvests from real replies. This
fixture serves two forks of a chain over MockRpcServer and runs a real
RpcSource against it, through registerReorgGuard and the rollback-depth
search.

It documents a gap it found along the way: the source answers both the reorg
comparison and the depth search from its own block cache, which still holds
the pre-reorg chain. The cached seam block reports the old hash, so the reorg
goes undetected; once detection is forced, the depth search confirms cached
blocks that no longer exist and stops at 101 where the fork is really at 99.
SourceManager.onReorg is what drops the cache, and Rollback.rollback calls it
only after getLastKnownValidBlock has run. Both cases here drop the cache
explicitly and assert the correct results, so the fixture pins the intended
behaviour rather than the current ordering.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RrYmidxrL75hDhP24nXmd6

* Invalidate source caches before the depth search, drop the unread block headers

Rollback.rollback dropped the sources' orphaned-chain state only after
getLastKnownValidBlock had run, so the depth search re-fetched the scanned
hashes through a cache the RPC source had filled on the abandoned fork. It
confirmed blocks that no longer existed and stopped short of the real fork,
leaving reorged data indexed. The call moves ahead of the search.

onReorg loses its rollback-target argument. The one implementation ignored it,
and it could not be otherwise: the deepest reorged block isn't known until the
search runs, and the search reads back through the very state the callback
clears — so pruning relative to a target would keep exactly the entries that
make it answer wrong. It now means "drop all of it".

EventItemsResponse.blocks carried a block header per returned block across napi
on every page for both EVM and Fuel. Nothing has read them since the block
store took over materialisation and reorg detection, so the field, the
BlockHeader/block DTOs, and the per-page header vector are gone.

Solana's threshold goes to 200 slots, far above the 32 blocks Tower BFT bounds
a fork by.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RrYmidxrL75hDhP24nXmd6

* Detect RPC reorgs from the range's own responses, not the block cache

The seam block (fromBlock - 1) is the only block a range shares with what the
store already scanned, so it is where a reorg at the boundary has to show up.
The source read it directly — and that read always hit the block cache, since
the seam is the previous range's toBlock, so the hash it compared was always
the one taken from the chain that range saw. A fork at the boundary matched
itself and went undetected.

It now reads fromBlock, which no earlier range touched, and takes the seam's
hash from its parentHash. Every hash reaching detection is then from this
range's own responses: the logs' blockHash, and the fetched block's hash and
parentHash. Same number of eth_getBlockByNumber calls — one per range, not one
cached plus one live.

The rollback fixture no longer has to drop the cache by hand to see the reorg,
and its second case moves to a fork the search reaches only through cached
blocks, so it still pins why onReorg runs before the depth search.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RrYmidxrL75hDhP24nXmd6

* Keep the validated hashes a rollback proved, drop only the forked ones

BlockStore.rollback dropped every hash above the restored progress block.
Two of those hashes were not stale: the depth search had just re-fetched and
validated the range up to the rollback target, and a chain rolled back for
cross-chain ordering never had a fork at all. Losing them means the refetch
of that range has nothing to compare against, so a source answering from a
lagging or orphaned fork goes undetected.

Rollback now takes the block above which hashes are suspect. The reorg chain
passes its validated target; every other chain passes null and keeps all of
them. Rows between the new progress and that bound survive as hash-only, the
same shape prune already leaves behind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Read the range's own blocks for detection, not the block cache

The seam and head hashes a range contributes to reorg detection were read
through the source's block loader. It caches by block number and is dropped
only on a reorg, so retrying a range after a transient failure answered both
reads from what the failed attempt saw — the fork the retry exists to reveal.
A range whose reorged blocks carry no matching logs then has nothing left to
disagree with, and the reorg goes unnoticed.

Both reads now load the block themselves and publish the result back to the
cache, so an event payload in the same block still costs one request, and a
single-block range reuses the seam read for its head.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Fold the two recoverable-retry helpers into one, escalate a stalled source

The behind-head and inconsistent-response retries were copies differing only
in backoff, log level and message, and each of their four call sites carried
the same trailing onReorg cache drop. One helper now takes the condition, and
the cache drop moved inside it — ahead of the backoff, so a sibling query on
the source stops reading orphaned-chain blocks immediately rather than for as
long as the wait lasts.

An inconsistent response is a reorg mid-request and clears on a refetch. One
that repeats is an endpoint serving blocks and logs from different chains, and
the chain stops progressing; past ~10 minutes of retries that is now an error
naming the source, not a warning repeated every minute forever.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Read the batch's hash snapshot in one call, seek to each gap

Building the snapshot cost one napi call per hashed block in the reorg
threshold — up to maxReorgDepth crossings per chain per batch, each taking the
store lock and allocating a hex string. One call now returns the numbers and
hashes as aligned columns.

The gap checkpoints then re-scanned the whole snapshot for every block
transition in the batch. The snapshot is ascending, so seek to the gap and
walk only its slice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Keep the timings of a failed getItems request

The native failure envelope carries the timings of the requests an operation
made before failing, and getBlockHashes records them. On the getItems path the
envelope was decoded, mapped to RateLimited or SourceBehindHead, and the
timings dropped — so exactly the requests that fail under throttling or head
drift were the ones missing from envio_source_request_*.

Both exceptions now carry the stats through to the retry, where SourceManager
records them like any other request.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Resume from the most recent checkpoint hash instead of refusing to start

Restoring reorg checkpoints treated a block number carrying two hashes as
fatal. Nothing in the schema forbids it — envio_checkpoints is keyed by id
alone — and a rollback interrupted between deleting the old checkpoints and
writing the new ones leaves exactly that behind. The throw then fails not just
that start but every one after it, with no way out but editing the table.

The restore now orders by checkpoint id and keeps the most recent hash per
block, which is what the registry it replaced did. Detection re-converges on
the next response.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Reject a block hash narrower than 32 bytes instead of zero-extending it

The store left-padded any hash up to 32 bytes into its comparison key. A
provider returning a truncated hash was therefore stored as a legitimate key,
and the correct hash for that block later read as a mismatch — surfacing a
malformed response as a reorg and a rollback rather than as the error it is.

Width is now part of the validation. The short markers test fixtures used are
widened where they enter a mocked response, which is where that accommodation
belongs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Share the block-hash paginator, fold the duplicated hex and conflict helpers

The EVM and SVM block-hash queries were the same ~60-line loop twice: range
derivation, guards, timed request with its stat, the error envelope, the
no-progress check, append, terminate. They had already drifted on the overlap
anchor. One driver now owns the seam and no-progress invariants; each source
supplies its query and, for SVM, its cursor-coverage bookkeeping.

Alongside: `merge` and `append_page` built the same mismatch by hand, and
`lowest_conflict` re-implemented the ordering `record_conflict` already does.
The crate's three hex parsers, which disagreed on prefixes, become one module
with the three contracts they actually needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Share the HyperSync block-hash wrapper, decode the failure envelope by schema

The EVM and SVM sources carried the same twelve-line wrapper around their
client's block-hash query; it belongs next to the failure mapping it uses.

The native failure envelope was unpacked with five levels of hand-rolled
JSON.Decode, silently dropping any stat entry that did not match. It is written
by our own serializer and has a fixed shape, so a schema states it once and
fails loudly instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Return the conflicting cells from the mismatch scan

The scan found the differing hashes and then threw them away, leaving the
caller to look both up again by key.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Simplify: keep the two retry helpers, drop the bisection and the page hook

Several of the earlier changes bought little for what they cost to read:

- The behind-head and inconsistent-response retries were folded into one
  function behind a variant and a positional triple. The duplication they
  shared was small and the names were the useful part; splitting them back out
  keeps the onReorg fix (inside the inconsistent one, ahead of the backoff) and
  the stall escalation, without the ladder.
- The gap checkpoints got a hand-rolled binary search over a list bounded by
  maxReorgDepth. The scan it replaced costs microseconds; the straight range
  filter is back.
- The block-hash paginator took a second closure only SVM used. The SVM fetch
  already knows the range it covered, so it marks its own coverage and the
  driver takes one closure.
- Restoring checkpoints deduped with a set and a reverse loop where a dict
  keyed by block number says the same thing.
- LazyLoader.set repeated the loader's eviction bookkeeping instead of sharing
  it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PvZPzzfer6g8BsUi9heweD

* Anchor block-hash pagination to the page that produced it

A page returning no rows kept the previous page's last block as the
overlap anchor. On SVM that anchor sits behind the cursor whenever a
window is all skipped slots, so the next request replayed the same empty
window, the cursor never advanced, and the source-behind-head error
reproduced on every retry.

The anchor is now the block the page itself ended on, and a rewound
request that covers no new ground drops the anchor and resumes from the
cursor instead of being reported as a stalled source — an overlap page
capped at the block it re-requested is this loop's own doing, not an
instance behind the head.

The SVM caller no longer marks coverage for a range the cursor did not
advance past, so that case surfaces through the paginator's own error
rather than a malformed coverage range.

Covers the paginator with unit tests over an injected backend: the
overlap seam, the empty page, the rewound page, a genuinely stalled
source, and an empty range.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JscwnEcKUbmcQbPtotGZ48

* Seek to each checkpoint gap instead of rescanning the hash snapshot

The snapshot is ascending, so the gap's first block is a bisection away.
Scanning it whole for every gap made checkpoint building quadratic in the
number of blocks a wide batch spans.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JscwnEcKUbmcQbPtotGZ48

* Address PR review comments on block-store reorg tracking

- block_store merge: reject cross-ecosystem pages at runtime (was debug-only),
  reset stale old-fork hashes above the divergence on detect-only reorgs, and
  prune hash-only observations below the reorg threshold so long no-event
  ranges stay bounded.
- EVM/SVM sources: keep a hash-only row for every returned block/slot header
  whose logs/instructions were all dropped by client-side routing, so a fork
  on such a block can still be detected.
- SourceManager: compute missing hashes only when there is no response
  conflict; replace em-dashes with hyphens in the reorg runtime files.
- e2e template test: give the outer `pnpm test` case a timeout cushion over
  the command timeout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Address more PR review comments on block-store reorg tracking

- block_store detect-only reset: clear only the stale hash field above the
  divergence instead of dropping the rows, so full block data other partitions
  buffered still materialises (regression from the previous reset).
- SourceManager: reduce the inconsistent-response stall threshold from ~10 to
  ~5 minutes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Address PR review comments: rollback semantics, block ref, cache revert

- Revert "Read the range's own blocks for detection, not the block cache": the
  block cache read is fine, since the next block comes from the RPC and its
  parentHash carries the reorg detection. Drops `LazyLoader.set` with it.
- Rollback now drops every block above its target, hashes included, on every
  chain. A non-reorg chain therefore replays without its hashes until it
  refetches, so its checkpoints in that range carry no hash.
- Move the source-cache drop into `getLastKnownValidBlock`, so the depth search
  cannot be run against a cache filled on the orphaned fork.
- FetchState: replace the `blockRef` record with a plain `latestFetchedBlock:
  int`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Bound the hash prune by processed progress, add SVM behind-head signal

- merge takes `pruneHashesBelow`, so a hash for a block the chain has not
  processed yet survives even when it falls outside the reorg threshold. A
  backfill sits far below the threshold, and those blocks are still to be read.
- SVM get_event_items reports a replica that has not reached the queried range
  through `source_behind_head_err`, like the paginator and the EVM/Fuel item
  paths already do, and the source re-raises the recoverable markers instead of
  burying them in a generic retry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Drop the detect-only reset, strip checkpoint dedup, unify retry backoff

- Detect-only reorgs no longer reset anything: the page overwrites what it
  observed and later pages converge the rest.
- Drop the reorg-checkpoint dedup and its ORDER BY. One block number carrying
  two hashes needs a rollback torn between its delete and its insert, and both
  run in the same transaction.
- One backoff schedule for every same-request retry, replacing the two linear
  ones and the exponential one.
- The block-hash paginator reports a behind-head instance from the overlap
  request instead of spending a second request from the cursor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Report a hash for every item block in the mock source

A real source returns the header of every block a matched log came from, but
the mock only reported the range's seam and its last block. Reorg detection
therefore never saw the blocks events actually landed on: checkpoints written
on them carried no hash, and the rollback depth search skipped them.

The rollback tests now search those blocks too, so they answer for them - with
a differing hash where the block belongs to the reorg they set up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

* Trim the block store on batch progress only

Merging a page no longer prunes. The progress-time prune already drops every
processed row and reduces the ones still inside the reorg depth to their hash,
and a response with no items still advances progress to the fetched frontier,
so it runs on empty ranges too. Dropping the merge-time prune removes the need
to bound it by the processed progress, and the rollback-guard rows it could
never recognise are trimmed by position like any other processed row.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aj6SS9KbG9mytdzuYnMs3a

---------

Co-authored-by: Claude <noreply@anthropic.com>

This branch was previously deployed

1 inactive deployment
internal — 9460ebce Deployed Jul 14, 2026 by DZakh via authorize #2169
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant