Skip to content

docs(ci): measure the integration collection split (no CI win; reverted) - #1004

Merged
mforce merged 3 commits into
mainfrom
perf/863-split-integration-collection
Oct 1, 2026
Merged

mforce merged 3 commits into
mainfrom
perf/863-split-integration-collection

Conversation

@mforce

@mforce mforce commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Why

docs/plans/863-integration-collection-split/01-measurement.md's 2026-09-30 amendment ruled out splitting the shared integration collection, but #1002 and #1003 removed both of its stated objections, and three clean local runs of main showed the shared collection running alone for 79-93 seconds at the end of every run. This PR tested the resulting hypothesis: split the collection into two halves, each with its own factory and database.

What this PR actually ships

No code change. The split was implemented, verified for correctness, measured locally (a clean, non-overlapping 47.3-second / 14.6% win), then measured on CI (ten samples per side, re-running only the Tests (integration) job). CI showed no measurable win: a mean difference of 2.4 seconds against a 27-second standard error, about 0.09 standard errors from zero. The split code is reverted in this branch; the net diff against main is one new file.

Scope

  • docs/plans/863-integration-collection-split/05-collection-split.md: the full method, both halves' class lists, the halves-alone correctness check (487/487 and 548/548, each completely alone in a fresh database), the vacuous-pass audit (every Assert.All/Assert.Empty/.CountAsync()/IgnoreQueryFilters() call site across all 94 classes checked against the fix(test): repair demo seed false pass and measure collection floor #1000 failure shape; none needed strengthening), build fingerprints proving each local run measured the intended code, and the local and CI measurements in full.

Verification

  • Both halves passed 100% alone before the code was reverted.
  • Local: 3 clean baseline runs (312.6-331.6 s) vs 3 clean split runs (273.2-277.7 s), non-overlapping, median saving 47.3 s.
  • CI: 10 samples per side. Split 405-586 s (median 518.0, mean 513.1). Baseline 433-571 s (median 545.5, mean 515.5). Ranges overlap almost completely; the split-side samples show a visible bimodal pattern (five runs at 405-483 s, five at 553-586 s) consistent with GitHub-hosted runner variance dominating over the code change.
  • The doc records a plausible mechanism for the local/CI gap: the local host has 12 cores, letting two ~153-second halves genuinely run concurrently; a standard ubuntu-26.04 runner has far fewer, so two collections may just compete for the same scarce cores instead of gaining real parallelism.

Refs #863, a broader issue about the integration suite's wall-clock floor. This PR answers one hypothesis from that issue with a measured no; the parent issue's disposition and whether another approach is worth trying are the owner's call.

mforce added 2 commits October 1, 2026 01:29
#1002 (one Postgres server) made an extra collection cost a database
clone instead of a container start. #1003 (demo-seed serialization)
stopped the heaviest nonshared work contending with itself. Both
objections that killed the original #863 split attempt are gone, and
three clean local runs on main (ad258d8) show the shared collection
running alone for 79-93 seconds at the end of every run.

IntegrationCollectionA and IntegrationCollectionB each get their own
ICollectionFixture<CluckworkWebApplicationFactory>, so each half gets
its own factory and migrated database and the two run concurrently
with each other. The 94 classes are split by measured test time
(152.7s / 152.7s) via greedy longest-processing-time packing, moving
whole classes only. The three default-account classes stay together
in half A; the two global-purge classes stay together in half B.
Both halves pass 100% alone in a fresh database (487/487 and
548/548), proving neither needs data from the other.

tools/test-timing/summarize.py now reports both halves separately
(shared_collection_halves) as well as their combined total
(shared_collection, unchanged shape).

See docs/plans/863-integration-collection-split/05-collection-split.md
for the method, both halves' class lists, and measurements.
Method, both halves' class lists, the halves-alone correctness check,
the vacuous-pass audit against #1000's failure shape, build
fingerprints proving each run measured the intended code, and three
clean interleaved local runs per configuration. CI section follows
once those runs land.
Ten CI samples per side (re-running only the Tests (integration) job,
never the whole workflow) show a mean difference of 2.4 seconds
against a 27-second standard error between split and baseline, about
0.09 standard errors from zero. The clean, non-overlapping 47.3-second
local win does not appear on CI, most likely because a standard
GitHub-hosted ubuntu-26.04 runner has far fewer cores than the
12-core host the local measurement ran on, so two collections
competing for the same scarce cores gain little from no longer being
one serial block.

The correctness work and both measurements stay recorded in
05-collection-split.md: both halves pass completely alone, the
vacuous-pass audit found nothing needing a fix, and the local win is
real. The code is reverted because the question this spike exists to
answer, whether the split lowers the wall time CI reports on every
PR, has a measured no for an answer.
@mforce mforce changed the title perf(test): split the shared integration collection into two halves docs(ci): measure the integration collection split (no CI win; reverted) Oct 1, 2026
@mforce
mforce merged commit 6166890 into main Oct 1, 2026
15 checks passed
@mforce
mforce deleted the perf/863-split-integration-collection branch October 1, 2026 04:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant