Skip to content

chore(ci): refresh the Test Core shard-timings dataset - #21826

Merged
objectstack-fleet[bot] merged 1 commit into
mainfrom
claude/shard-timings-refresh-37262126122
Oct 6, 2026
Merged

objectstack-fleet[bot] merged 1 commit into
mainfrom
claude/shard-timings-refresh-37262126122

Conversation

@github-actions

@github-actions github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Refreshes scripts/test-shard-timings.json, the balancing input for the Test Core
shard split. Opened automatically by .github/workflows/shard-timings-refresh.yml.
Every byte came out of scripts/measure-test-shard-timings.mjs; nothing here was
hand-edited, and no bound, timeout or matrix entry was touched.

Source

Measured across 1 accumulated run(s) of the HOURLY schedule run of CI on
main — the full-battery run (#16467). A push run on main is affected-only and
is not a measurement of the workspace, so no push run feeds this file.

No single green run measures the whole workspace either — turbo's cache is namespaced
per shard and only main pushes write it, so a
package whose inputs have not changed is a HIT and the generator refuses hits rather
than recording a replay as a duration. Runs are therefore accumulated, each fenced by
its own --run group, until every package the committed dataset holds is measured
again; a package seen in several of them gets the median of those observations.

  • https://github.com/objectstack-ai/objectstack/actions/runs/37262126122

  • Newest run in the set: 37262126122, commit 75ddcd1b41e74823b7bb9f3fe9e159a189fa2a91 — the date this refresh carries.

  • Every run above had all six Test Core (N/6) jobs conclude success with its six
    run-summary artifacts still retained; runs that were cancelled, failed or had lost
    their artifacts were rejected by name in the log before any of these were used.

All 72 package weights were measured in these runs; nothing was carried.

⚠️ This PR references #16173 and #16222 but does NOT carry a closing keyword for them,
because a weekly lane cannot know which cards a given run ought to retire. If this is
the first refresh to land, retire those two by hand as part of merging it.

Measured per-shard suite time on the newest run in the set

1: 759s  |  2: 2100s  |  3: 844s  |  4: 647s  |  5: 802s  |  6: 796s

Predicted bins, before and after

BEFORE  partition-test-shards: self-test OK (72 measured packages -> 72 shard items, 6 shards, max/mean 1.12x <= 1.3x, floor 1391s, bins 1391/1208/1206/1208/1208/1208s, file-level slices: none)
AFTER   partition-test-shards: self-test OK (72 measured packages -> 72 shard items, 6 shards, max/mean 1.04x <= 1.3x, floor 1703s, bins 1703/1615/1615/1615/1617/1616s, file-level slices: none)

No checks will start on this PR by themselves

It was opened with the Actions GITHUB_TOKEN, and GitHub's recursion guard means a
PR opened that way triggers no workflow runs. Push any commit to the branch, or close
and reopen the PR, to start CI.

Refs #16464, #16173, #16222.

Regenerated by .github/workflows/shard-timings-refresh.yml from the test-core-run-summary artifacts of 1 accumulated run(s) (37262126122), newest 37262126122 at 75ddcd1. Generated, never hand-edited.
@objectstack-fleet

Copy link
Copy Markdown
Contributor

Stand-down note from the domain:devx seat 1 · session_01VDtqoecgES7ScQYGbFVDRv · 2026-10-06T00:54Z.

The merge queue dropped this PR (CI_FAILURE). The failure is not this PR's.

  • Failing check: Lint & Type Check → job Lint & Repo Gates → step PM dispatch-gates self-test, in merge-group run 37393301718. The step reported "1 of 1976 case(s) failed".
  • Why it is not this PR's:
    • The same step fails in main's own scheduled full run 37394652870, at be97cf3c93, which does not contain this PR.
    • The hourly card hourly full run: red on main (Lint & Type Check) #21924 (p1) tracks that red on main.
    • This PR changes only scripts/test-shard-timings.json. Every PR-stage check passed on head 18a0dad5bb.
  • Fix: none exists yet; hourly full run: red on main (Lint & Type Check) #21924 is open. No change is ported here, because the cause is outside this diff.

This PR goes back into the queue once main's full run is green again.


Generated by Claude Code

@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 6, 2026
Merged via the queue into main with commit f2aa0c9 Oct 6, 2026
32 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/shard-timings-refresh-37262126122 branch October 6, 2026 04:22
akarma-synetal pushed a commit to akarma-synetal/framework that referenced this pull request Oct 7, 2026
…mings dataset (objectstack-ai#21966)

Fixes objectstack-ai#21758
Clause-②: no

This PR covers triage's direction `5982304306`, steps 2 to 4. Step 1
(option A, the re-measure) is the refresh that already landed as
`f2aa0c9fad`. The only thing left open is the seat's reading on the next
hourly full run after this lands.

The diff is one comment block in `scripts/partition-test-shards.mjs`. No
code changed. `FILE_SHARDED_PACKAGES` is unchanged (still empty), and so
are `test-shard-timings.json`, `ci.yml` and every timeout.

## Step 2: the docblock is corrected

The `FILE_SHARDED_PACKAGES` docblock argued that the CLI "fits whole
until ~1852s". It reached that by solving the bound against a 733.33 s
entry. That entry was the CLI's two slice windows summed within one run
(run `36380128221`, the only sample that refresh had). It was not a
whole-suite window. Once the CLI ran whole, its windows read 1659.03 s
and 1667.97 s, which is 2.26 to 2.27 times that entry. The conclusion
held, because both readings are under ~1852 s, but the figure behind it
was wrong.

The block now records that, and it re-derives the bound on the dataset
that measured the CLI whole.

## Step 3: option B is not indicated (the arithmetic)

**Dataset.** `scripts/test-shard-timings.json` at `origin/main`
`01e0f71a`, blob `12c2460c00`. That blob is identical to `f2aa0c9fad`,
the objectstack-ai#21826 refresh.
- It was measured from run `37262126122`: 72 packages, 9781.33 s in
total.
- `@objectstack/cli` reads 1702.69 s and is now the heaviest item.
- `@objectstack/dogfood`, the only `--exclude` in ci.yml, is not in the
dataset.

**The rule used is the partitioner's own, not a paraphrase.**
`meetsBound` (pins 2 and 3) requires both of these:
- the heaviest bin is at most `MAX_SHARD_OVER_MEAN` (1.3) times the
mean;
- no single item is heavier than 1.3 times the mean.

When the CLI is the heaviest item and sits alone in a bin, both halves
reduce to C ≤ (1.3/6)(R + C), where C is the CLI's weight and R is
everything else.

**Numbers.**
- R = 8078.64 s.
- Mean = 9781.33 / 6 = 1630.22 s.
- Bound = 1.3 × 1630.22 = 2119.29 s.
- CLI/mean = 1702.69 / 1630.22 = **1.044**, inside 1.3.
- Solving for C: C_max = 1.3 × 8078.64 / 4.7 ≈ **2234.5 s**. That is
1.31 times the CLI's dataset entry, and 1.27 times its worst reading
since the refresh (1753.66 s, run `37415122516`).

| CLI weight | max/mean | heaviest item |
|---|--:|---|
| whole, as in the dataset (1702.69 s) | 1.044 | 1703 s (cli) |
| sliced at 2 | 1.002 | 1135 s (spec) |
| whole at its worst since (1753.66 s) | 1.070 | 1754 s (cli) |

Slicing at 2 would lower the maximum, but the bound already holds at n =
1. Pin 3c's own derivation refuses a `{ '@objectstack/cli': 2 }` entry
on this dataset ("at 1 the split already meets 1.3x (max/mean 1.04x,
heaviest item 1703s against a 1630s mean). Retire the entry"). So
**option B is not taken**.

Because `FILE_SHARDED_PACKAGES` did not change, the Test Core matrix and
the required-check names (`Test Core (N/6)` and the aggregate `Test
Core`) do not change either.

## The 6 bins the landed dataset produces (full list, as the hourly run
splits it)

| shard | predicted | share of mean | items | contents |
|--:|--:|--:|--:|---|
| 1/6 | 1702.69 s | 1.044 | 1 | cli |
| 2/6 | 1614.77 s | 0.991 | 12 | spec, service-messaging, plugin-audit,
mcp, … |
| 3/6 | 1615.33 s | 0.991 | 15 | metadata-protocol, service-analytics,
verify, cloud-connection, … |
| 4/6 | 1615.40 s | 0.991 | 15 | objectql, service-automation, client,
lint, … |
| 5/6 | 1616.73 s | 0.992 | 15 | rest, plugin-auth, plugin-approvals,
driver-turso, … |
| 6/6 | 1616.41 s | 0.992 | 14 | runtime, plugin-security, driver-sql,
plugin-sharing, … |

The partitioner's own self-test prints the same bins: `bins
1703/1615/1615/1615/1617/1616s`.

## Step 4: second samples (reported here, not written into the dataset)

**Source.** Each reading is a turbo execution window for a task that
actually ran (a cache MISS). It comes from the `report-test-timings`
merged table that the aggregate `Test Core` job echoes into its log,
read with the GitHub MCP `get_job_logs` read tool.
- The `test-core-run-summary-*` artifacts and the raw job-log REST
endpoint are both unreadable from this container: `gh api
…/actions/jobs/112121376726/logs` refuses the redirect to
`productionresultssa16.blob.core.windows.net`.
- The merged table lists only the 10 slowest packages per run.

Each ratio is the reading divided by the landed dataset value.

| package (dataset) | 37421524959 schedule @ 76fec88 | 37415122516
push @ 3c7785d | 37419396793 push @ 76fec88 | 37416453417 schedule
@ 3c7785d |
|---|--:|--:|--:|--:|
| plugin-auth (364.98 s) | 393.33 s, 1.08× | 387.53 s, 1.06× | 231.44 s,
0.63× | cache replay, NOT MEASURED |
| metadata-protocol (566.45 s) | 372.89 s, 0.66× | 549.06 s, 0.97× | not
in the top-10 (cutoff 150.51 s), NOT MEASURED | cache replay, NOT
MEASURED |
| cli (1702.69 s) | 1723.05 s, 1.01× | 1753.66 s, 1.03× | 1721.19 s,
1.01× | 1033.74 s, 0.61× |

Seat 1's pre-refresh executed readings (`5982793172`), divided by the
new values:
- plugin-auth: 297.82 / 417.15 / 320.88 s, which is 0.82× / 1.14× /
0.88×;
- metadata-protocol: 511.20 / 426.17 s, which is 0.90× / 0.75×.

Every sample for both packages is under 1.5× against the landed dataset.
The highest are plugin-auth at 1.14× and metadata-protocol at 0.97×.

## Early reading toward the acceptance (not the seat's post-landing
reading)

This PR changes no code, so the partition on `main` is already the one
above. Two post-refresh hourly full runs exist:
- **Shard 1/6** (cli alone): measured/predicted is 0.61× in
`37416453417` and 1.01× in `37421524959`.
- **Shards 2 to 6: NOT MEASURED here.** A shard's ratio needs every
package window on it. Those windows are only in the run-summary
artifacts, which this container cannot reach, and the merged table names
10 packages. In `37421524959` all ten listed packages read between 0.66×
and 1.21×.

## Acceptance notes (noted, not filed)

- **Shard 1/6 wall time.** Shard 1 is the whole CLI alone. Its job took
31m37s (`37421524959`), 34m39s (`37415122516`), 32m09s (`37419396793`)
and 20m42s (`37416453417`). The current timeout is a temporary 45
minutes, and its written revert condition returns it to 30. This PR
changes no timeout (fenced), and the 1.3× ratio bound holds. Handed on:
objectstack-ai#16465 and the owner of that revert condition.
- **ci.yml comments are now stale** (ci.yml is objectstack-ai#16465's to edit, so they
were not touched here):
- the drift-step block still cites the CLI as "458.15s recorded,
1231.52s measured";
- "WHY SIX" names `@objectstack/spec` as the heaviest indivisible suite,
but on this dataset the CLI is.
  - Handed on: objectstack-ai#16465.
- **The refreshed dataset is again a single run.** Its `provenance.runs`
is `["37262126122"]`. The CLI alone on shard 1 read 1033.74 s and
1723.05 s one hour apart. Handed on: the refresh lane's next weekly run.
- **spec's cost on the new partition is unknown.** spec moved from a bin
of its own into a 12-package bin. The weights are contended wall clock
(ci.yml "WHY SIX"), and spec was a cache replay in all four post-refresh
runs read here. Its window under the new placement is NOT MEASURED. If
shard 2 drifts, look there first.

## Verification (head `cb3f093c`)

- `node scripts/pm/dispatch-gates.mjs --commands --repo
objectstack-ai/objectstack` derived 30 commands from the merge base
`01e0f71ad`, the same 30 the dispatch named.
  - All 30 exited 0, with each exit code written to disk.
- `--ran` reconciliation: `✓ dispatch-gates --ran: 30 derived famil(ies)
accounted for — 30 run, 0 NOT-MEASURED (a DERIVED zero …)`.
- `node scripts/partition-test-shards.mjs --self-test`: `self-test OK
(72 measured packages -> 72 shard items, 6 shards, max/mean 1.04x …,
floor 1703s, bins 1703/1615/1615/1615/1617/1616s, file-level slices:
none)`, where the elided part states the 1.3x bound.
- `pnpm check:pm-dispatch-gates`: `✓ dispatch-gates self-test: 1976
cases pass.`
- `pnpm check:nul-bytes`: `OK (scanned 10322 text file(s) … no raw ASCII
control bytes)`.
- Extra runs, for rosters under `scripts/`: `node
scripts/check-published-list-mirrors.mjs` and `node
scripts/measure-test-shard-timings.mjs --self-test` both exit 0.
- No `*.test.*` file names this script (zero hits; the control grep for
`scripts/` over the same pathspec has 383 hits).
- **eslint, narrowed to this one file:**
- `eslint --no-inline-config --format json
scripts/partition-test-shards.mjs` gives 1 file, 0 errors, 0 warnings.
- The file is in the population of `eslint.config.mjs`'s
`**/*.{…,mjs,…}` object; `--print-config` shows 2 rules and
parserOptions `{ecmaVersion, sourceType}`.
- The config enables no type-aware linting (no `parserOptions.project`
or `projectService` anywhere), so a comment edit in one file cannot
change any other file's result. The full `pnpm lint` run is CI's.
- No changeset: root `scripts/` only, private, nothing publishes, so
`skip-changeset`.

The holder of this card ran out of session tokens. This PR continues the
existing branch `claude/issue-21758-shard-fit-after-refresh` as a
takeover (`Release:`/`Claim:` `6010756437`), by `domain:devx` seat 2,
session `session_01VF48aw8RPG6wzDnMgp6rtw`.

---
_Generated by [Claude
Code](https://claude.ai/code/session_01VF48aw8RPG6wzDnMgp6rtw)_

Co-authored-by: Claude <noreply@anthropic.com>
akarma-synetal pushed a commit to akarma-synetal/framework that referenced this pull request Oct 7, 2026
…red/predicted, warning past 1.3x (objectstack-ai#21998)

Fixes objectstack-ai#16465
Clause-②: no

Wires the Test Core shard timing-drift check the tree has carried
unwired since objectstack-ai#16173, adds the card's 1.3x as a warning tier under its
red, and restates the `ci.yml` comments that the wiring and the objectstack-ai#21487 /
objectstack-ai#21826 changes left stale. Root `scripts/` and one workflow, nothing
published: `skip-changeset`.

## The ruling this follows

Triage `5925054826`, as the unlock `6013064225` restates it, verbatim:

- "one red rule, the in-tree `partition-test-shards.mjs --check-drift`
at `MAX_MEASURED_OVER_PREDICTED = 1.5`;"
- "the card's 1.3× is a `::warning::` only;"
- "⛔ no second red rule, so no 70%-of-`timeout-minutes` red."

The maintainer's authority on the card: 「同意你的建议,你负责执行派发所有可行的优化」.

## What changed

**`.github/workflows/ci.yml`, the Test Core shard job only**

- New step `Check this shard's timing drift`: `--check-drift` over
`.turbo/runs/*.json` with `--label "Test Core (N/6)"` (the matrix
shard), with no `if:` and no `continue-on-error`. It sits between the
run-summary upload and the completeness guard, above the attestation
pair, so a drift red also withholds that shard's attestation. A shard
with no packages writes no summary; the step reports that as NOT
MEASURED and exits 0, because the script itself treats zero inputs as a
usage error.
- The "⛔ THE DRIFT STEP IS DELIBERATELY NOT WIRED HERE YET" block is
replaced by a description of what the step does. The stale CLI figures
("458.15s recorded, 1231.52s measured") and "the CLI is halved across
two runners" are gone. The comment now cites the dataset as objectstack-ai#21826
refreshed it (provenance run 37262126122, CLI 1702.69s) and the second
sample below.
- **WHY SIX** now names `@objectstack/cli` (1702.69s) as the heaviest
indivisible suite, ahead of `@objectstack/spec` (1134.86s). The
6/7/8/10-shard table is re-derived on the current dataset with the
partitioner's own `partition()` / `balanceOf()`: 1.04x / 1.22x / 1.39x /
1.74x, with the max at 1703s at every count. The old claim that the
self-test "pins this arithmetic" is corrected: what it pins is pin 3,
the heaviest package within 1.3x of the mean at `SHARD_COUNT`.
- **The 45-minute timeout comment.** ⛔ The value stays 45. The revert
condition is now stated against measured wall time instead of "once
objectstack-ai#16173 lands": back to 30 only when the slowest shard's job wall time
stays at or under 24 minutes (80% of 30) on every scheduled run for a
week. On the 16 main runs after the objectstack-ai#21826 refresh, shard 1/6 (the CLI
alone) read 6m14s to 35m43s of job wall time, and was over 30 minutes on
9 of them (for example 34m39s in run 37415122516 and 35m43s in run
37453598388). Two present-tense "30-minute wall" phrases in the same job
now say "`timeout-minutes` wall".
- **Slice-leg prose.** The build-closure step and the test step now say
that `FILE_SHARDED_PACKAGES` has been empty since objectstack-ai#21487, so no shard
runs a slice today.

**`scripts/partition-test-shards.mjs`**

- `WARN_MEASURED_OVER_PREDICTED = 1.3` sits beside
`MAX_MEASURED_OVER_PREDICTED = 1.5`, which is unchanged. It is the same
ratio over the same executed windows; it is not `MAX_SHARD_OVER_MEAN`,
and its docblock says why the two are kept separate.
- `driftReport()` gains `warned`, the band strictly between the two
bounds and exclusive of the red, so one shard gets one verdict.
- `renderDriftVerdict()` returns the four verdicts (OK, WARN, DRIFT, NOT
MEASURED) without printing them. WARN emits exactly one `::warning
title=Test Core shard timing drift::…` line on stdout and exits 0; DRIFT
exits 1, with the same remedy text as before.
`escapeWorkflowCommandMessage()` keeps the message on one line.
- New self-test battery `drift warning tier under the red (objectstack-ai#16465)`,
with 12 cases, registered in `SELF_TEST_BATTERIES` (floor 12). The
roster floor goes from 10 to 11. The cases cover:
- both edges of the band, the red's exclusivity and a healthy 1.18x
reading;
  - replay-only input;
- band non-empty (`WARN` below the red) and not below the balance
tolerance;
- the rendered annotation line, the red's exit code, and command
escaping.

Nothing in the self-test prints a workflow command
(`check:self-test-workflow-commands` green).

Not touched: the attestation schema, `scripts/test-shard-timings.json`,
`measure-test-shard-timings.mjs`, and the required-check names. `Test
Core` and `Test Core (N/6)` are unchanged, and `check:required-contexts`
is green.

## The second sample (carried item 2): executed windows only

The `test-core-run-summary-*` artifacts could not be read from this
container: the egress policy denies
`productionresultssa17.blob.core.windows.net`. Job logs reach only their
last 5000 lines through `get_job_logs`. So the sample comes from each
run's `Test Core` timing table, which the aggregator echoes into its
log. The table carries turbo's own execution windows through
`samplesFromSummary()`, which is the same reader and the same exclusions
as the gate. Replayed packages are listed there and excluded here.

| shard | reading | source runs |
|---|---|---|
| 1/6 (CLI alone) | **exact**, 0.61x, 1.05x, 0.65x, 0.78x, 1.02x, 1.03x,
1.01x, 0.79x | the 8 scheduled runs 37416453417, 37427862594,
37433381795, 37440129185, 37446949708, 37453598388, 37460325624,
37467795959 |
| 2/6 to 6/6 | **NOT MEASURED whole.** The table prints only the 10
slowest packages, so a shard's lighter members have no window here |
same runs |
| (small-set push runs) | exact on every shard that ran: 0.90x (spec
alone), 0.48x (rest alone), 0.67x (create-objectstack alone); 0.18x
(client alone) | 37422473456; 37413386379 |

On shards 2 to 6, the measured heavy members and the turbo
`--concurrency=4` capacity bound (the sum of a shard's task windows is
at most 4x its test step) give these results:

- **Proven under 1.5x:** 8 shard-run pairs (for example shard 3 in
37427862594, at most 0.82x).
- **Undecided** everywhere else. No measured package read 1.5x or more
in any of the 9 near-full runs.
- **Highest package reading:** `@objectstack/spec`, 1.45x / 1.39x /
1.43x on the three runs that executed it (37433381795, 37453598388,
37467882762).
- **Nearest the bound: shard 2/6.** spec carries 70% of its prediction,
so its 11 unmeasured members would have to average 1.61x to 1.77x before
the shard reached 1.5x.

**Stop condition:** not met. No executed-window reading put any shard at
or over 1.5x, so the red is wired as ruled. The per-shard numbers, the
bounds and this PR's own run are on objectstack-ai#16465 in the `os-dev-report`.

## Verification (at `430859c2f`)

- Derived gates: `node scripts/pm/dispatch-gates.mjs --repo
objectstack-ai/objectstack --ran …` printed "59 derived famil(ies)
accounted for — 59 run, 0 NOT-MEASURED (a DERIVED zero — all 59 recorded
an exit code and none of them is 3)". Every gate exited 0.
- Four of them (`check:dts-closure`, `check:dual-build-cjs-loads`,
`check:lean-entry-closure`, `check:sourcemap-no-sources-content`) first
answered exit 3 (PREREQUISITE NOT MET). They exited 0 once a workspace
build had run through the verify lock: 73 tasks, `VERDICT command-exit
0`.
- `node scripts/partition-test-shards.mjs --self-test` → `self-test OK
(72 measured packages -> 72 shard items, 6 shards, max/mean 1.04x …)`.
- `--check-drift` on synthetic summaries: 1.06x gives OK, exit 0; 1.34x
gives WARN plus one `::warning` line, exit 0; 1.54x gives DRIFT, exit 1.
A replay sits beside each, and they are excluded.
- The step's bash, run against an empty directory, gave NOT MEASURED and
exit 0. Against one summary it gave the verdict.
- `check-governed-merges --test` on the final two paths: "0 of 2 path(s)
hit the register … NOT governed", with 494 changed lines.

## Acceptance notes

- **Shards 2 to 6 still lack a whole-shard reading from main.** This
step is what produces one. The first scheduled runs after landing print
it per shard, and the WARN annotations show where.
- **This PR's own Test Core run is a thin sample.**
- Its affected set, simulated with `turbo ls --affected` and the
cross-package union at the merge base, is `@objectstack/spec`,
`@objectstack/client` and `@objectstack/driver-sql`.
- spec's `test` leg is expected to replay, which makes it NOT MEASURED.
client and driver-sql run alone.
  - So it exercises the wiring rather than the dataset.
- **Expect the warning on shard 2/6 of cold full runs.** spec read 1.39x
to 1.45x against its single-run 1134.86s entry. That is the warning
doing its job, and the remedy is the refresh lane, not this step.
- **Small-sample hazard, observed and not filed.** `driftReport()` has
no floor on predicted seconds. A PR whose shard executes only a
few-second package gets a ratio of that package alone. Every
lone-package reading found here was under 1.0x, so this is noted, not
filed.
- **The 24-minute (80%) revert criterion is this PR's choice**, stated
with its reason in the comment. The current 45 is already at 79% on
shard 1/6's slowest reading (35m43s).

---
_Generated by [Claude
Code](https://claude.ai/code/session_01VF48aw8RPG6wzDnMgp6rtw)_

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/m skip-changeset PR has no user-facing published change; bypasses the changeset gate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants