Repository navigation
docs(ci): measure the integration collection split (no CI win; reverted) - #1004
Merged
Merged
Conversation
#1002 (one Postgres server) made an extra collection cost a database clone instead of a container start. #1003 (demo-seed serialization) stopped the heaviest nonshared work contending with itself. Both objections that killed the original #863 split attempt are gone, and three clean local runs on main (ad258d8) show the shared collection running alone for 79-93 seconds at the end of every run. IntegrationCollectionA and IntegrationCollectionB each get their own ICollectionFixture<CluckworkWebApplicationFactory>, so each half gets its own factory and migrated database and the two run concurrently with each other. The 94 classes are split by measured test time (152.7s / 152.7s) via greedy longest-processing-time packing, moving whole classes only. The three default-account classes stay together in half A; the two global-purge classes stay together in half B. Both halves pass 100% alone in a fresh database (487/487 and 548/548), proving neither needs data from the other. tools/test-timing/summarize.py now reports both halves separately (shared_collection_halves) as well as their combined total (shared_collection, unchanged shape). See docs/plans/863-integration-collection-split/05-collection-split.md for the method, both halves' class lists, and measurements.
Method, both halves' class lists, the halves-alone correctness check, the vacuous-pass audit against #1000's failure shape, build fingerprints proving each run measured the intended code, and three clean interleaved local runs per configuration. CI section follows once those runs land.
Ten CI samples per side (re-running only the Tests (integration) job, never the whole workflow) show a mean difference of 2.4 seconds against a 27-second standard error between split and baseline, about 0.09 standard errors from zero. The clean, non-overlapping 47.3-second local win does not appear on CI, most likely because a standard GitHub-hosted ubuntu-26.04 runner has far fewer cores than the 12-core host the local measurement ran on, so two collections competing for the same scarce cores gain little from no longer being one serial block. The correctness work and both measurements stay recorded in 05-collection-split.md: both halves pass completely alone, the vacuous-pass audit found nothing needing a fix, and the local win is real. The code is reverted because the question this spike exists to answer, whether the split lowers the wall time CI reports on every PR, has a measured no for an answer.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
docs/plans/863-integration-collection-split/01-measurement.md's 2026-09-30 amendment ruled out splitting the sharedintegrationcollection, but #1002 and #1003 removed both of its stated objections, and three clean local runs ofmainshowed the shared collection running alone for 79-93 seconds at the end of every run. This PR tested the resulting hypothesis: split the collection into two halves, each with its own factory and database.What this PR actually ships
No code change. The split was implemented, verified for correctness, measured locally (a clean, non-overlapping 47.3-second / 14.6% win), then measured on CI (ten samples per side, re-running only the
Tests (integration)job). CI showed no measurable win: a mean difference of 2.4 seconds against a 27-second standard error, about 0.09 standard errors from zero. The split code is reverted in this branch; the net diff againstmainis one new file.Scope
docs/plans/863-integration-collection-split/05-collection-split.md: the full method, both halves' class lists, the halves-alone correctness check (487/487 and 548/548, each completely alone in a fresh database), the vacuous-pass audit (everyAssert.All/Assert.Empty/.CountAsync()/IgnoreQueryFilters()call site across all 94 classes checked against the fix(test): repair demo seed false pass and measure collection floor #1000 failure shape; none needed strengthening), build fingerprints proving each local run measured the intended code, and the local and CI measurements in full.Verification
ubuntu-26.04runner has far fewer, so two collections may just compete for the same scarce cores instead of gaining real parallelism.Refs #863, a broader issue about the integration suite's wall-clock floor. This PR answers one hypothesis from that issue with a measured no; the parent issue's disposition and whether another approach is worth trying are the owner's call.