Repository navigation
Regression harness compares results across different package environments (Manifest drift), producing spurious regressions #338
Description
Activity
@matt-pharr (and @logan-nc) wanted to flag this to your attention since you're more familiar with the regression harness. I came across this issue while chasing down changes in the regression tests for wayyyyyy too long only for it to end up being a Manifest/package difference and had Claude generate this summary for me.
I'm thinking the first option seems reasonable (i.e. pass a warning if the Manifest is different), since I think its reasonable for the regression to assume all packages are the same. Any thoughts on this?
Hmmmm... I dislike the tracking of the manifest due to worries that it will pin us unnecessarily in the past when we should be following updates of dependencies as much as possible to keep up-to-date. It's better to stay up to date than to find ourselves stale and full of update conflicts 5 years down the line. So we don't want to force everyone to a lowest-common denominator of conservative package pinning.
I suppose the easiest fix is just to make the default regression test the "force" case where it runs both the new and the old in the users local env rather than using shared pinned values.
Next step I could see would be to do option (1) of hashing the manifest of the shared pin values and only forcing if there is a change. Ideally, we'd have this provide instructions to update the local env if it's behind on some package updates. We would then only update the git pins if the local env is all greater than or equal as far as the manifest goes (i.e. keep the pins pinned to whoever has the most up-to-date local env). I'd rather have a situation that is pulling everyone up to the most recent package set rather than slowing everyone down to stale ones.
Neither of the above feels like a complete / clean vision... but hopefully the general principle makes sense? Thoughts @matt-pharr ?
Ah yup I just discovered this independently today.We clearly need to be at least resolving dependency changes better. This is the worry I expressed during the original hackathon and IIRC Brendan Lyons' response was to track the manifest.
One option is we could track the manifest and then use dependabot to open a monthly pull request of consolidated updates, on which we can run tests/regression harness, and review.
Reacted by Nikolas Logan- marked Regression harness compares refs against local using different dependency environments #346 as a duplicate of this issue
on Jul 31, 2026 ok. that sounds good. As. long as it doesn't fail quietly and let us get stale too easily 😅
@logan-nc I am getting a set of package updates now that flip the sign of the torque in the d3d ideal example and change b_res by 25%... So I think requiring we pay attention to dependency bumps is a good thing
Reacted by Nikolas LoganYikes!
Adding one more failure mode I hit independently yesterday (originally clauded as #353, which I'm closing as a duplicate). It's adjacent to the original manifest issue, but this is about which code becomes the baseline, not which packages.
From claude:
The harness silently uses a stale local baseline.
resolve_ref(regression-harness/src/utils.jl:22-36) resolves a branch name with a plaingit rev-parse --verify --quiet <ref>, i.e. the local branch pointer, falling back toorigin/<ref>only when the local ref doesn't exist. Nothing compares the resolved commit to its upstream tracking branch. My localdevelopwas 120 commits behindorigin/develop, so--refs develop,localwas benchmarking against three-week-old code and nothing in the report said so.The runner does log the commit's author date (
runner.jl:241,runner.jl:385), but an author date is easy to read past and doesn't answer the question that matters — is this the current baseline?This compounds the Manifest issue rather than duplicating it: a stale baseline commit is also the one most likely to have a badly mismatched package set. In my case the two effects together produced a hard failure rather than just wrong numbers —
developat41cfbcb3died withMethodError: no method matching Float64(::Vector{Float64})and reported every quantity as FAILED. That's the extractor meeting re-resolved packages whose APIs had moved, not a bug at that commit.Suggested addition to the acceptance criteria here: warn when a resolved ref is behind its remote tracking branch (
git rev-list --count <ref>..origin/<ref>), reporting the count, and say plainly that the baseline is not current.For the record, the Manifest half of #353 reproduced this issue exactly:
--cases diiid_n1 --refs develop,localreported 14 changed / 34 unchanged (resonantb^r9.80%, Chirikov 6.90%, island half-widths 2.21%, ODE steps 2033→2006, Mercier/ballooning checksums all flipped), while re-running the same two code states as two git refs —--refs develop,HEAD, so both sides instantiate identically — gave48 unchanged, every quantity bit-identical at0.0e+00. 153 of 388 packages differed between the two manifests, including OrdinaryDiffEqCore 3.33.1 vs 3.29.0 and FastInterpolations 0.4.18 vs 0.4.8.#395 adds committed, per-Julia-minor dependency pins under
ci/manifests/, which happen to supply the mechanism this issue's root causes 1 and 2 need. Noting it here so the two do not get solved twice.What #395 adds
ci/manifests/Manifest-v1.11.tomlandManifest-v1.12.toml— full resolves, tracked in git (they do not collide with theManifest.tomlentry in.gitignore).ci/manifests/update.jl— regenerates a pin for whichever Julia runs it, resolving a bare copy ofProject.tomlin a temp dir so a developer's ownManifest.tomlis untouched.ci/manifests/project.sha256—Project.tomlchecksum at generation time, so a stale pin can be detected without starting Julia.
It was added to stop CI re-resolving against the live registry on every run, but the artifact is exactly what the harness is missing.
Why this is a better fix than pinning both refs to one manifest
The controlled re-test in the description had to force the
developworktree onto the working tree'sManifest.tomlby hand. That isolates the source diff, but it is not what you want in general: if a commit legitimately changed[deps]or[compat], evaluating it under a different commit's dependency set measures the wrong thing.Because the pins are committed,
git worktree addat any commit brings that commit's pinned manifest with it. The environment travels with the commit, so each ref is measured under the dependency set it was actually written against, and identical environments collapse to identical results automatically — no manual pinning step, and no wrong answer when deps genuinely differ.Sketch
Root causes 1 and 2 —
create_worktree(regression-harness/src/utils.jl:74-83) copies the pin matching the running Julia into the worktree before instantiate:pin = joinpath(worktree_path, "ci", "manifests", "Manifest-v$(VERSION.major).$(VERSION.minor).toml") isfile(pin) && cp(pin, joinpath(worktree_path, "Manifest.toml"); force=true)
Guarded on
isfile, so commits predating #395 fall back to today's behaviour rather than erroring.Root cause 3 — the
runstable (regression-harness/src/database.jl:5-18) gains an environment column, andis_cachedkeys on it. The hash of the resolvedManifest.tomlis the natural fingerprint, andjulia_versionis already recorded inside each manifest. That still matters independently: it is what stops a baseline cached before this change from being compared against a run made after it.Caveats
- Only fixes refs at commits that carry pins. Baselines cached before CI - IMPROVEMENT - Pin the CI dependency set per Julia minor version #395 stay confounded until the environment key of root cause 3 lands and invalidates them, which is an argument for doing 3 alongside 1 and 2 rather than after.
- The pins track a resolve, not the newest one. They need regeneration when
[deps]/[compat]change —CLAUDE.mdnow carries that rule — so a harness run reflects the pinned set, not whatever is newest. For regression comparison that is the point, but it does mean a genuine upstream regression (cf. FastInterpolations 0.4.15 -> 0.4.17 changes b_res by 25% and flips the sign of the toroidal torque #347, FastInterpolations changingb_resby 25%) will not surface until the pins are refreshed. Worth a periodic refresh, and worth deciding deliberately rather than by accident.
Not proposing to do this inside #395, which is deliberately CI-only.
Closing. #361 implemented proposed fixes 1 and 2 and the stale-baseline failure mode raised in the comments:
- Every worktree in a comparison is pinned to the working tree's resolved Manifest.toml.
- Each run records its Julia version, host, Manifest hash and thread counts. Results are cached under that environment key, so
a result from a different environment is re-run instead of reused. - The report prints a banner when compared refs ran in different environments, or when a ref is behind its remote.
Fix 3 (a tracked environment for generating golden values) was deliberately deferred to the golden-values work in #228. The
PATH-julia hang and single-sample caching are tracked separately in #418.- added a commit that references this issue
on Oct 1, 2026
Regression harness compares results across different package environments (Manifest drift), producing spurious regressions
Summary
The regression harness caches numerical baselines keyed only on
(commit_hash, case_name), with no fingerprint of the package environment that produced them. BecauseManifest.tomlis untracked and each worktree run callsPkg.instantiate()against a bare checkout, a baseline cached weeks ago is silently resolved against older package versions than a freshlocalrun. The resulting machine-epsilon (~1e-16) differences in library math (BLAS/SciML/ArrayInterface) are then amplified by the adaptive ODE step controller and by ill-conditioned near-resonant diagnostics into large apparent regressions — up to 200% — that are pure environment artifacts, not code changes.This matters because the harness is mandated on every PR (per
CLAUDE.md). A reviewer can waste significant time chasing a phantom, or — worse — "fix" correct code to mask an artifact.Root cause (confirmed in code)
Three independent factors combine:
Manifest.tomlis untracked. A worktree created for a cached commit has no pinned Manifest.regression-harness/src/utils.jl:74-83(create_worktree) does a baregit worktree add --detach.Each run resolves the environment fresh. With no Manifest present,
Pkg.instantiate()picks the newest versions compatible withProject.toml's[compat]at run time.regression-harness/src/runner.jl:43-54(RUNNER_SCRIPT_TEMPLATE) →%INSTANTIATE%expands toPkg.instantiate()(default; only skipped with--no-instantiate).regression-harness/src/runner.jl:409-423against--project=$worktree_path.The cache has no environment key. Results are stored/looked up by
(commit_hash, case_name)only.regression-harness/src/database.jl:5-18(runstable,UNIQUE(commit_hash, case_name)) — no Julia version, no Manifest hash, no package-set hash.is_cached(regression-harness/src/database.jl:63-67) ignores the environment entirely.Net effect:
regress --refs develop,localcompares a develop baseline produced under whatever package set was newest at cache-population time against a fresh local run under today's package set.Evidence
Investigating a branch that only refactors
FourierTransforms/Vacuum (no changes toEulerLagrange.jl/Riccati.jl/Fourfit.jlor the equilibrium metric), the harness reported fordevelop(cached) vslocal:Controlled re-test — same source diff, but the develop worktree pinned to the identical
Manifest.tomlas the working tree (source code the only variable), 4 threads throughout:The develop worktree instantiated without pinning had resolved newer versions (e.g.
4.6.1→4.7.0,7.25.0→7.28.1in the SciML/ArrayInterface stack). Pinning the Manifest collapsed every reported regression to machine epsilon. Two back-to-back runs of the same environment are bit-identical, so this is not thread-schedule noise.Mechanism: a 1e-16 change in equilibrium/EL-matrix math (from a different library version) crosses an adaptive-step acceptance boundary (384↔385 steps), which reshuffles the integration ψ-grid, and the near-resonant singular-coupling quantities (Chirikov, Δ′, torque) are ill-conditioned enough to turn that seed into double-digit-percent swings.
Proposed fixes (increasing effort)
Fingerprint the environment and warn on mismatch (minimum). Store a hash of the resolved
Manifest.toml(or Julia version + package-set) alongside each cached run, and print a loud warning (or auto-invalidate) when a cached ref's fingerprint differs from the current run's. Cheapest change; converts a silent confound into a visible one. Requires arunsschema column + a check inis_cached/reporter.Run both refs in one shared, pinned environment at compare time. Resolve a single Manifest once and reuse it for every worktree (
--projectpointed at a shared depot, or copy the working tree's Manifest into each worktree beforeinstantiate). Makes source code the only variable by construction.Track
Manifest.toml(or a harness-specific pinned Manifest) so worktree checkouts are reproducible, and instantiate against it.Workaround (today)
Regenerate the baseline in the current environment before trusting a report:
--forcere-runs the develop ref under the same packages aslocal, eliminating the cross-environment artifact.