You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
test: simlin-serve watcher test runtime scales with fseventsd health, can bust the pre-commit cap on passing tests #945
The 16 simlin-serve watcher integration tests (tests/integration/watcher_smoke.rs, watcher_merge.rs, watcher_git.rs) each perform multiple sequential real-filesystem watch round-trips through notify-debouncer-full's RecommendedWatcher (FSEvents on macOS). Their wall-clock runtime is therefore a linear function of the host's fsevents notification latency -- an external daemon's health, not anything in the code under test.
Evidence
On a machine with a degraded fseventsd daemon (pinned at ~100% CPU, delivering events with ~0.9-1.3s latency instead of the normal <50ms -- measured with an independent node fs.watch probe):
The 16 watcher tests take 155s (~10s each, ~0% CPU -- pure waiting on event delivery).
The whole workspace cargo test goes from a healthy <90s to 235s+, busting the pre-commit hook's 180s wall-clock cap (scripts/pre-commit line 160) and blocking all commits on that machine -- even though every test passes.
On a healthy machine the same suite fits comfortably inside the cap.
Why it matters
A gate whose pass/fail depends on the health of an unrelated system daemon is fragile: the developer's code is fine, every test is green, and commits are still blocked (with --no-verify prohibited by CLAUDE.md).
scripts/pre-commit (the 180s cap and its exit-124 diagnostic)
Possible approaches (not prescriptive, not mutually exclusive)
Batch multiple assertions per watch session to cut the number of sequential round-trips per test -- reduces the multiplier on per-event latency without changing what is exercised.
Per-test event-latency telemetry, so a cap trip can report "the watcher tests spent N seconds waiting on fsevents delivery" instead of asserting "this is genuine test runtime" -- turning a misdiagnosis into an actionable message even if the runtime itself is not fixed.
tech-debt.md entry 37 -- deterministic macOS event-loss for pre-existing files; correctness, not runtime.
Discovery
Identified during conveyor-engine branch work (unrelated to that branch's changes) on a machine whose fseventsd was degraded, where every commit was blocked by the cap despite an all-green test run.
Problem
The 16
simlin-servewatcher integration tests (tests/integration/watcher_smoke.rs,watcher_merge.rs,watcher_git.rs) each perform multiple sequential real-filesystem watch round-trips throughnotify-debouncer-full'sRecommendedWatcher(FSEvents on macOS). Their wall-clock runtime is therefore a linear function of the host's fsevents notification latency -- an external daemon's health, not anything in the code under test.Evidence
On a machine with a degraded
fseventsddaemon (pinned at ~100% CPU, delivering events with ~0.9-1.3s latency instead of the normal <50ms -- measured with an independent nodefs.watchprobe):cargo testgoes from a healthy <90s to 235s+, busting the pre-commit hook's 180s wall-clock cap (scripts/pre-commitline 160) and blocking all commits on that machine -- even though every test passes.Why it matters
--no-verifyprohibited by CLAUDE.md).scripts/pre-commitlines 168-172) says "Test binaries were already compiled before the timed run, so this is genuine test runtime" and points at per-test budgets. That message was written for the GH build: local pre-commit cargo test 180s cap is unachievable on macOS (~449s clean run, dominated by simlin-serve fsevents/e2e tests; cap also covers compile) #703 fix (compile time excluded from the cap) and is correct for CPU-bound regressions, but it misdiagnoses this failure mode -- the time is not test work, it is waiting onfseventsddelivery, and nothing in the output lets the developer see that.Components affected
src/simlin-serve/tests/integration/watcher_smoke.rssrc/simlin-serve/tests/integration/watcher_merge.rssrc/simlin-serve/tests/integration/watcher_git.rsscripts/pre-commit(the 180s cap and its exit-124 diagnostic)Possible approaches (not prescriptive, not mutually exclusive)
notify::PollWatcher), trading FSEvents fidelity for delivery independent of daemon health. Tradeoff: the tests no longer exercise the production FSEvents path (same tradeoff already noted in test: simlin-serve watcher integration tests are flaky on macOS (fixed wall-clock timeouts vs FSEvents latency) #916 approach 3 and tech-debt entry 37).Relationship to existing tracking (duplicates checked)
cargo test --no-runbefore the timed step. This issue is the remaining post-compile gap: even pure test execution can blow the cap when the external daemon is slow, and the diagnostic added for build: local pre-commit cargo test 180s cap is unachievable on macOS (~449s clean run, dominated by simlin-serve fsevents/e2e tests; cap also covers compile) #703 actively misattributes it.Discovery
Identified during conveyor-engine branch work (unrelated to that branch's changes) on a machine whose
fseventsdwas degraded, where every commit was blocked by the cap despite an all-green test run.