You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
perf gate cannot observe chain-traversal regressions: every fixture scenario is a one-patch chain #849
The v19 base/head performance gate exercises PatchDiscovery (materialization is that path) but cannot detect a chain-traversal regression, because no scenario ever builds a multi-patch chain.
Rewritten 2026-08-21. The original diagnosis said the corpus was below the checkpoint interval and recommended raising baseNodeCount. That was measured and falsified — see Falsified hypothesis below. The real cause is one level down, and the recommended fix and its constraints have changed accordingly.
Root cause
scripts/performance/PerformanceFixture.ts:118 writes the entire corpus segment inside a singleruntime.patch(...) call:
asyncfunctionappendCorpus(runtime,start,count,propertyBytesPerNode,seed){awaitruntime.patch((patch)=>{for(letoffset=0;offset<count;offset+=1){// ... addNode / setProperty / addEdge, all in ONE patch}});}
So baseNodeCount controls how much payload a patch carries, not how many patches exist. Every scenario materializes a chain of exactly one patch (incremental-materialize: two, base + suffix).
Measured confirmation — three runs, CI corpus (25 base / 5 suffix), local arm64:
Scenario
replayedPatches
gitCommandCount (median, MAD)
cold-materialize
1
2758, MAD 0
warm-materialize
0
543, MAD 0
incremental-materialize
1
2788, MAD 0
replayedPatches: 1 is the direct evidence: a 25-node corpus replays one patch.
Falsified hypothesis
The original issue claimed 25 base nodes never reach DEFAULT_CHECKPOINT_INTERVAL = 64 (src/domain/warp/CheckpointPolicy.ts:7), so raising the corpus above 64 would make incremental-materialize a genuine short-tail case. That was tested directly:
Run A — 25 base / 5 suffix: cold 2758 cmds incremental 2788 cmds
Run B — 80 base / 5 suffix: cold 5956 cmds incremental 5794 cmds
At 80 nodes — comfortably above the interval — incremental still costs the same as cold. Raising the node count does not create a checkpoint, because the interval counts patches and the fixture writes one. The counts scale with payload volume, which is the corpus dimension baseNodeCount actually controls.
Why it matters
incremental-materialize is designed to be the short-tail-behind-a-checkpoint case, and it is the only scenario that could discriminate a bounded read from an unbounded one. With a one-patch chain there is no traversal to bound, no checkpoint to respect, and no difference to measure.
Concretely: PR #847 fixed a read that spawned one git show per patch commit, serially. On this fixture that is one spawn. The gate would have shown the regression as green at any corpus size. The same PR then shipped, and self-review caught, a read that ignored stopAtSha and indexed the entire first-parent history — also invisible here for the same reason.
Suggested fix
Give the fixture a patch-count dimension, independent of node count — e.g. a patchCount spec that splits appendCorpus across N runtime.patch(...) calls. Set it above DEFAULT_CHECKPOINT_INTERVAL for the base segment so a checkpoint genuinely forms and incremental-materialize measures a short tail behind it.
Once the chain is deep, three metrics discriminate at once: gitCommandCount (structural, MAD 0 — see below), cpuTotalMedianMs, and peakHeapUsedBytes for the unbounded-index variant.
Constraint: this shifts every measured baseline, so benchmarks/v19/policy.json and calibration.json need regenerating on the reference runner (ubuntu-24.04, 4 CPU, per calibration.json.environment). Hand-written thresholds from a developer machine would not be trustworthy. Recalibration has to happen in CI. This is the part that genuinely must wait.
What has landed
Two pieces of protection are in place; neither closes this issue.
Structural assertions in the unit suite (perf: batch patch-chain and git-cas object reads #847): PatchDiscovery.batched.test.ts asserts the bulk read is issued with stopAt and that only the tail is emitted (bulkRecordsEmitted). Deterministic, cannot flake, verified to fail against the unbounded read.
A Git-command gate in the perf policy (perf: batch patch-chain and git-cas object reads #847): absolute.gitCommandMedian and relative.gitCommandRegressionRatio, with the counts reported in the gate summary. Counts are machine-independent and measured at MAD 0, so this needed no reference recalibration to land. Verified end-to-end against real measured data: a synthetic 10x spawn regression fails both checks while CPU reads +0.0%.
The gate is the right instrument, and it is now wired and calibrated. But on a one-patch chain it gates object and payload traffic rather than traversal depth. It becomes the sharp tool for this issue only once the fixture above exists.
changed the title [-]perf gate cannot observe unbounded chain reads: corpus is below the checkpoint interval[/-][+]perf gate cannot observe chain-traversal regressions: every fixture scenario is a one-patch chain[/+]on Aug 21, 2026
The
v19 base/head performancegate exercisesPatchDiscovery(materialization is that path) but cannot detect a chain-traversal regression, because no scenario ever builds a multi-patch chain.Root cause
scripts/performance/PerformanceFixture.ts:118writes the entire corpus segment inside a singleruntime.patch(...)call:So
baseNodeCountcontrols how much payload a patch carries, not how many patches exist. Every scenario materializes a chain of exactly one patch (incremental-materialize: two, base + suffix).Measured confirmation — three runs, CI corpus (25 base / 5 suffix), local arm64:
replayedPatchesgitCommandCount(median, MAD)replayedPatches: 1is the direct evidence: a 25-node corpus replays one patch.Falsified hypothesis
The original issue claimed 25 base nodes never reach
DEFAULT_CHECKPOINT_INTERVAL = 64(src/domain/warp/CheckpointPolicy.ts:7), so raising the corpus above 64 would makeincremental-materializea genuine short-tail case. That was tested directly:At 80 nodes — comfortably above the interval —
incrementalstill costs the same ascold. Raising the node count does not create a checkpoint, because the interval counts patches and the fixture writes one. The counts scale with payload volume, which is the corpus dimensionbaseNodeCountactually controls.Why it matters
incremental-materializeis designed to be the short-tail-behind-a-checkpoint case, and it is the only scenario that could discriminate a bounded read from an unbounded one. With a one-patch chain there is no traversal to bound, no checkpoint to respect, and no difference to measure.Concretely: PR #847 fixed a read that spawned one
git showper patch commit, serially. On this fixture that is one spawn. The gate would have shown the regression as green at any corpus size. The same PR then shipped, and self-review caught, a read that ignoredstopAtShaand indexed the entire first-parent history — also invisible here for the same reason.Suggested fix
Give the fixture a patch-count dimension, independent of node count — e.g. a
patchCountspec that splitsappendCorpusacross Nruntime.patch(...)calls. Set it aboveDEFAULT_CHECKPOINT_INTERVALfor the base segment so a checkpoint genuinely forms andincremental-materializemeasures a short tail behind it.Once the chain is deep, three metrics discriminate at once:
gitCommandCount(structural, MAD 0 — see below),cpuTotalMedianMs, andpeakHeapUsedBytesfor the unbounded-index variant.Constraint: this shifts every measured baseline, so
benchmarks/v19/policy.jsonandcalibration.jsonneed regenerating on the reference runner (ubuntu-24.04, 4 CPU, percalibration.json.environment). Hand-written thresholds from a developer machine would not be trustworthy. Recalibration has to happen in CI. This is the part that genuinely must wait.What has landed
Two pieces of protection are in place; neither closes this issue.
Structural assertions in the unit suite (perf: batch patch-chain and git-cas object reads #847):
PatchDiscovery.batched.test.tsasserts the bulk read is issued withstopAtand that only the tail is emitted (bulkRecordsEmitted). Deterministic, cannot flake, verified to fail against the unbounded read.A Git-command gate in the perf policy (perf: batch patch-chain and git-cas object reads #847):
absolute.gitCommandMedianandrelative.gitCommandRegressionRatio, with the counts reported in the gate summary. Counts are machine-independent and measured at MAD 0, so this needed no reference recalibration to land. Verified end-to-end against real measured data: a synthetic 10x spawn regression fails both checks while CPU reads+0.0%.The gate is the right instrument, and it is now wired and calibrated. But on a one-patch chain it gates object and payload traffic rather than traversal depth. It becomes the sharp tool for this issue only once the fixture above exists.