Skip to content

perf(ci): lower the integration suite floor — split the shared integration collection #863

Description

@mforce

AMENDED 2026-10-01 — re-scoped again; supersedes the 2026-09-30 amendment
below.
That amendment ruled out splitting the integration collection and
said to reopen it "only if a later measurement shows a durable shared-only
tail". That condition is now met, locally.

Two merged changes removed both objections to the split:

Measured on three clean local runs of current main (ad258d8f; 245 classes,
94 in integration), from the full per-class timing file: the shared
collection ends at 316.9–320.4 s of a 317.6–321.0 s suite; the last
substantive nonshared class (DemoSeedAttributionTests, the end of the
demo-seed lane) ends at 227.8–241.1 s. The shared collection runs alone
for 79–93 s at the end of every run.
AppDbContextDesignTimeFactoryTests
finishes later but with ~0 s of work — it is a DisableParallelization
collection — so it is not the floor.

New scope: split the shared integration collection, starting at width
two
, within the correctness boundary in §D6 of
docs/plans/863-integration-collection-split/01-measurement.md.
Projection, not measurement: two balanced halves of ~308 s of shared work
would end around 165–170 s each, leaving the demo-seed lane (~230–240 s) as
the new floor — a suite near 240 s, roughly 25% below today's ~320 s. Width
three buys nothing past that floor unless the seed lane also shrinks. Two
halves running side by side also add CPU load, so the real figure may land
higher; measure it.

CI must be measured directly. CI's integration time swings ~140 s on
identical code — larger than any change here — so compare ~10 CI runs per
side. The earlier finding that CI's two sides finish within ~23 s of each other
predates #1002 and #1003 and must be re-established, not assumed.

Shipped under this issue so far: #1000 (repaired a DemoSeedTests false
pass), #1002 (~11.5% local), #1003 (~6.9% local).


AMENDED 2026-09-30 — re-scoped. The premise below is measured false; do not
plan from it.
#863 was "split the serialized integration collection". A
measurement pass (PR #1000,
docs/plans/863-integration-collection-split/01-measurement.md, and
this comment)
established that the serialized collection is no longer the critical path,
so splitting it at any width now buys 0–23.5 s, about 4%. The body's claim
that the other 142 classes "fit inside the block's shadow" held on a 12-core
host; on current CI the nonshared side ends after the shared side, or within
23.5 s of it. The original 331 s → 90 s estimate has no support on current main.

The real floor is a handful of very slow nonshared classes, and two of them
are single-test classes each running a full demo seed:
DemoSeedDisabledOwnerTests 190.9 s for one test and
DemoSeedAttributionTests 182.8 s for one test — 373.7 s of the 1,450.9 s
nonshared work in two tests. Then SeedCommandTests 151.9 s / 8,
ProcessRoleGuardTests 81.4 s, OtlpSubprocessExporterTests 69.5 s / 11,
AccountScopedIdentityMigrationTests 59.0 s, OneShotVerbMinimalConfigTests
58.6 s. These launch the app as a subprocess or create databases. A class
cannot be split
, so the tallest of them is the suite's floor no matter what
happens to collections.

New scope: lower the cost of those classes. Share or cache the demo-seed
work the two single-test classes each pay in full; reduce redundant subprocess
launches and per-class database creation. Measure first, exactly as this issue
already required — the lesson that killed the original premise applies again.

Out of scope, unchanged: test deletion (#775 decision 1) and higher worker
counts. Also out of scope now: splitting the integration collection. Reopen
that only if a later measurement shows a durable shared-only tail; the
correctness boundary for doing so safely is recorded in §D6 of the measurement
document and should be re-read rather than re-derived.

Already shipped by the measurement pass: a real false pass in DemoSeedTests
(asserted only IsSuccess, so it passed in 0.25 s without seeding — 18.55 s of
actual work) and database-wide count coupling in AccountProvisioningTests.


Split out of #839, which merged as a variance fix (#861). This is the lever that moves the floor.

The problem

dotnet test overlaps collections but runs tests within a collection serially. The shared integration collection is one collection holding 95 classes and 1,024 of 1,815 tests — 316.1s of test work that cannot overlap itself.

That 316s serial block is the suite's wall-clock floor. Measured on the #861 run: 331.5s total wall clock, of which the shared collection occupied 316.1s end to end. #861 stops the block landing at the end of the timeline (the 434s outlier); it cannot shrink the block.

Everything else — 142 classes, 852.4s of test work — already runs in parallel across 12 cores and fits inside the block's shadow. Splitting the shared collection is the only change that lowers the floor.

First task: measure, do not split

#839's own lesson applies. Its stated premise (per-class container startup) was already solved before anyone looked. Before designing a split, establish:

  1. What state the 95 classes actually share. The default account, base reference data (grades, roles, packed-unit conversions), Simulation:* config, and any advisory-lock assumptions. The shared fixture exists because they share something; find out what, per class, before assuming two classes are separable.
  2. Which classes are already isolated in practice — those that never write, or write only rows nobody else reads. Those are the candidates.
  3. What the floor would become at 2, 3, and 4 collections of the resulting groups. A split that produces two 158s blocks halves the floor; one that produces 300s + 16s does nothing.

Constraints

Explicitly not this issue

Test deletion and higher worker counts. #775 decision 1 argues against deletion; #776 landed coverage to answer "what is untested", which is never "what is redundant".

Known remaining cost, for whoever sizes this

The parallel work is lumpy: six classes hold 374s of the 852s. OtlpSubprocessExporterTests 80.5s for 11 tests, ProcessRoleGuardTests 71.7s, SeedCommandTests 61.0s, OneShotVerbMinimalConfigTests 58.7s. These repeatedly launch the app as a subprocess. Lowering the collection floor below roughly that height makes this the next floor, so measure both before committing to a split width.

Refs #839

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions