Repository navigation
test(web): enforce a precache ceiling and a Dashboard script-duration tripwire - #933
Merged
Merged
Conversation
… tripwire (#825) verify-sw.mjs now sums the precache manifest's real file sizes and fails above a 1,900 KiB ceiling, derived from the actual measured baseline (1,806.38 KiB at 75388d2) plus the two known pending costs (#828, #835), with a behavior-level node:test self-test that mutation-proves the guard fails closed. A new Playwright spec, tools/simulation/ui/specs/dashboard-script-duration.spec.ts, measures Dashboard's script execution time via CDP Performance.getMetrics against the real seeded stack and fails past a wide regression tripwire, printing the measurement every run; both were run against a live isolated stack and torn down. Supersedes the issue's 2026-09-16 "measure and report, no gate" decision, per this slice's brief; docs/designs/822-mui-revamp.md's D9 and §7 are amended in place to record why and to correct an arithmetic error an independent review caught (the projected post-#835 total is ~1,907 KiB, over this ceiling by ~7 KiB, which D7.2's already-documented h1/h2 fallback is sized to cover). Closes #825.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
Implements #825's scope: an explicit precache ceiling, enforced in CI, plus the script-duration signal #674's runtime-baseline record recommended alongside it. Closes #825.
web/scripts/verify-sw.mjssums the precache manifest's real file sizes (byte-accurate, readingdist/the same way it already extracts the manifest) and fails above a 1,900 KiB ceiling. Derivation is in the code comment and indocs/designs/822-mui-revamp.md's D9: measured baseline at75388d2is 1,806.38 KiB; web: retire NumberField and the hand-rolled tooltip positioning #828 (~-18 KiB) and web: give display text an optical size using Inter's opsz axis #835 (+118.9 KiB, already measured) are the two known pending costs, netting a projected ~1,907 KiB — over this ceiling by ~7 KiB. That's not smoothed over: D7.2 already names a fallback (one display-size cut forh1/h2, worth 24.1 KiB) for exactly this case, so the ceiling stays a real forcing function rather than being inflated to whatever the next projection is.web/scripts/verify-sw.test.mjsis a newnode:testbehavior-level test (subprocess against a crafteddist/, following the same self-test pattern as the two npm-audit gates inci.yml), wired into thewebCI job. Mutation-proven: lowering the ceiling in a throwaway copy makes it fail with the expected message.tools/simulation/ui/specs/dashboard-script-duration.spec.tsis a new Playwright spec, picked up automatically by the existing PR-gatinge2e-smoke.yml(no new CI wiring needed — it sharestestDir: ./specs). It signs in as Owner, client-side-navigates to Dashboard 5 times, and reads CDPPerformance.getMetricsdeltas (same domainsrc/ax.tsalready uses) to measure script execution time, printing the median every run and failing past a wide 600ms tripwire. This is a regression tripwire, not a tight budget, on purpose: CI runner CPU is shared and variable, so a tight number would be exactly the "wrong guard that reads as safety" this repo's own AGENTS.md warns against for every other guard — precache bytes are environment-independent and get the tight ceiling; timing gets the wide one.docs/designs/822-mui-revamp.mdD9 and §7 are amended in place (struck through, not deleted) to record two things: this supersedes the issue's 2026-09-16 "measure and report, no gate" decision (explained below), and the corrected arithmetic an independent review caught before this PR was opened.Why this supersedes the "no gate" decision
The issue's last comment (owner, 2026-09-16) deferred gating until #826, #827, #828, and #835 had all landed and the number had a trend. Only #826 and #827 have landed; #828 and #835 are still open. This PR's brief explicitly asked for CI enforcement of both ceilings now, so it proceeds on that instruction rather than waiting, and says so plainly in the D9 amendment instead of silently re-deciding it. The chosen 1,900 KiB ceiling is derived to already account for both pending changes, so it should not need revisiting the moment they land.
Component plan
Row 3 of the sequence table (
docs/designs/822-mui-revamp.md, §"one component library, fourteen slices"):#825budget | D9;verify-sw.mjsextension; the "adds weight without retiring code" rule | none. This is groundwork — CI tooling and a measurement spec, not a screen — and the table's own screen-impact column already says "none" for this slice.Verification
npm run build && npm run verify:sw(web/): passes, prints[precache budget] 1806.38 KiB / 1900 KiB ceiling (93.62 KiB headroom).node --test scripts/verify-sw.test.mjs(web/): 2/2 pass; mutation-tested by lowering the ceiling to 100 KiB in a throwaway copy, confirmed it fails with the expected over-ceiling message.npm run typecheckandnpm run test:coverage(web/): clean, 92.06%/88.94%/87.96%/94.76% stmt/branch/func/line coverage, no regressions.npm run typecheck(tools/simulation/ui/): clean.docker-compose.sim.ymlunder a throwawaycw825-script-timecompose project, distinct ports, torn down after — never the sharedcluckwork-simstack). It passed:runs=56,122,124,127,141ms medianMs=124.5. Mutation-tested by lowering its ceiling to 1ms against the same live stack, confirmed it fails with a real measured value.python3 -c "import yaml; yaml.safe_load(...)") and the newwebjob step placement double-checked against the existing self-test pattern.Independent review
Dispatched a read-only Paseo review (codex/gpt-5.6-sol,
auto-reviewmode — the closest to read-only Codex offers, plus an explicit no-edit instruction, since no pure read-only mode is available) against the complete diff, the issue, and the design doc, before this PR was opened. It returned 4 findings, all accepted and fixed:All four fixes were verified:
verify-sw.test.mjsreran green (2/2) with the corrected message text,npm run typecheckreran clean in bothweb/andtools/simulation/ui/, and the corrected precache math was independently confirmed (python3 -c "print(1806.38 - 18 + 118.9)"→ 1907.28).Remaining risk
The script-duration spec's 600ms tripwire is calibrated from exactly one real measurement plus a wide safety margin, not a distribution — CI runner variance is the reason it's wide rather than tight, but the first few real CI runs are the actual test of whether the margin holds. Watch
e2e-smoke.ymlon this PR and the next few merges; tighten only against a measured trend, per the spec's own comment.