Skip to content

perf gate cannot observe chain-traversal regressions: every fixture scenario is a one-patch chain #849

Description

@flyingrobots

The v19 base/head performance gate exercises PatchDiscovery (materialization is that path) but cannot detect a chain-traversal regression, because no scenario ever builds a multi-patch chain.

Rewritten 2026-08-21. The original diagnosis said the corpus was below the checkpoint interval and recommended raising baseNodeCount. That was measured and falsified — see Falsified hypothesis below. The real cause is one level down, and the recommended fix and its constraints have changed accordingly.

Root cause

scripts/performance/PerformanceFixture.ts:118 writes the entire corpus segment inside a single runtime.patch(...) call:

async function appendCorpus(runtime, start, count, propertyBytesPerNode, seed) {
  await runtime.patch((patch) => {
    for (let offset = 0; offset < count; offset += 1) {
      // ... addNode / setProperty / addEdge, all in ONE patch
    }
  });
}

So baseNodeCount controls how much payload a patch carries, not how many patches exist. Every scenario materializes a chain of exactly one patch (incremental-materialize: two, base + suffix).

Measured confirmation — three runs, CI corpus (25 base / 5 suffix), local arm64:

Scenario replayedPatches gitCommandCount (median, MAD)
cold-materialize 1 2758, MAD 0
warm-materialize 0 543, MAD 0
incremental-materialize 1 2788, MAD 0

replayedPatches: 1 is the direct evidence: a 25-node corpus replays one patch.

Falsified hypothesis

The original issue claimed 25 base nodes never reach DEFAULT_CHECKPOINT_INTERVAL = 64 (src/domain/warp/CheckpointPolicy.ts:7), so raising the corpus above 64 would make incremental-materialize a genuine short-tail case. That was tested directly:

Run A —  25 base / 5 suffix:  cold 2758 cmds   incremental 2788 cmds
Run B —  80 base / 5 suffix:  cold 5956 cmds   incremental 5794 cmds

At 80 nodes — comfortably above the interval — incremental still costs the same as cold. Raising the node count does not create a checkpoint, because the interval counts patches and the fixture writes one. The counts scale with payload volume, which is the corpus dimension baseNodeCount actually controls.

Why it matters

incremental-materialize is designed to be the short-tail-behind-a-checkpoint case, and it is the only scenario that could discriminate a bounded read from an unbounded one. With a one-patch chain there is no traversal to bound, no checkpoint to respect, and no difference to measure.

Concretely: PR #847 fixed a read that spawned one git show per patch commit, serially. On this fixture that is one spawn. The gate would have shown the regression as green at any corpus size. The same PR then shipped, and self-review caught, a read that ignored stopAtSha and indexed the entire first-parent history — also invisible here for the same reason.

Suggested fix

Give the fixture a patch-count dimension, independent of node count — e.g. a patchCount spec that splits appendCorpus across N runtime.patch(...) calls. Set it above DEFAULT_CHECKPOINT_INTERVAL for the base segment so a checkpoint genuinely forms and incremental-materialize measures a short tail behind it.

Once the chain is deep, three metrics discriminate at once: gitCommandCount (structural, MAD 0 — see below), cpuTotalMedianMs, and peakHeapUsedBytes for the unbounded-index variant.

Constraint: this shifts every measured baseline, so benchmarks/v19/policy.json and calibration.json need regenerating on the reference runner (ubuntu-24.04, 4 CPU, per calibration.json.environment). Hand-written thresholds from a developer machine would not be trustworthy. Recalibration has to happen in CI. This is the part that genuinely must wait.

What has landed

Two pieces of protection are in place; neither closes this issue.

  1. Structural assertions in the unit suite (perf: batch patch-chain and git-cas object reads #847): PatchDiscovery.batched.test.ts asserts the bulk read is issued with stopAt and that only the tail is emitted (bulkRecordsEmitted). Deterministic, cannot flake, verified to fail against the unbounded read.

  2. A Git-command gate in the perf policy (perf: batch patch-chain and git-cas object reads #847): absolute.gitCommandMedian and relative.gitCommandRegressionRatio, with the counts reported in the gate summary. Counts are machine-independent and measured at MAD 0, so this needed no reference recalibration to land. Verified end-to-end against real measured data: a synthetic 10x spawn regression fails both checks while CPU reads +0.0%.

    The gate is the right instrument, and it is now wired and calibrated. But on a one-patch chain it gates object and payload traffic rather than traversal depth. It becomes the sharp tool for this issue only once the fixture above exists.

Activity

  1. changed the title [-]perf gate cannot observe unbounded chain reads: corpus is below the checkpoint interval[/-] [+]perf gate cannot observe chain-traversal regressions: every fixture scenario is a one-patch chain[/+] on Aug 21, 2026
  2. added this to the v19.1.0 milestone on Aug 24, 2026
  3. added
    type:debtDebt, rot, or structural risk.
    priority:asapImmediate release pressure.
    status:activeSomeone is actively working this issue.
    area:testingPrimary work area: testing.
    on Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:testingPrimary work area: testing.priority:asapImmediate release pressure.status:activeSomeone is actively working this issue.type:debtDebt, rot, or structural risk.

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions