Skip to content

[Decision] Lint & Repo Gates aborts at the first failing gate — 120 sequential steps, no continue-on-error, so one red hides an unmeasured tail #12413

Description

@yinlianghui

Split out of #12173 by the domain:devx @ objectstack seat (#6023). That card carried two independent findings; its ① is dead (already fixed by #11934, verified at source — see #12173). This is ②, which is untouched and real. ⛔ This seat picks among none of the options.

Dispatch ruling on #12173 put ② out of the dev's scope precisely because it changes the repo's CI contract, so it was re-measured and reported, never acted on.

The finding

Lint & Repo Gates runs its gates sequentially in one job and aborts at the first non-zero exit. A single red is therefore not "one gate to fix" — it is one gate to fix plus an unknown number that never ran.

Re-derived on origin/main at b000ab59bb (the filing-day numbers were stale):

filing day measured now
named steps in the job — 120 (lint.yml:111-3111)
real continue-on-error keys inside the job — 0 (the single textual hit at :2451 is a comment)
gate steps after the specimen's failure point 46 10

The specimen: check:where-matcher now sits at :2768 and the ObjectQL double limit gate — the step that failed on the incident this was filed from — immediately follows at :2789.

⚠️ The 46 → 10 correction matters for pricing this, and it cuts against urgency: on that specific failure the unmeasured tail was 10 gates, not 46. The general shape stands; the worst case is bounded by where in the 120 the failure lands.

四维分析

① 实际业务需求. 拉力是真实但间歇的:每次红都留下一条未测量的尾巴,而 PR 作者读到的是「一个门禁要修」。⚠️ 但重新测出的 10(而非 46)说明单次事故的代价比立卡时以为的小,而代价随失败位置变化 —— 落在第 5 步和落在第 115 步完全是两回事。没有人测过失败位置的分布。

② 平台长远合理性. 这一维指向「一次运行报告全部失败」:一个门禁农场的价值在于它告诉你全部欠什么,而不是最先撞上的那一件。⚠️ 反方向也真实:continue-on-error 会让 job 在有真实失败时仍然报绿,除非另加一个汇总步骤 —— 那本身就是新机制,而且是削弱门禁方向的改动,按协议属人工地板。

③ 避免 AI 写代码犯错. 强烈指向拆分或全量报告:当前形状下,dev 修好第一个红、推送、再撞第二个红,一次事故要烧掉多个 CI 往返 —— 而每一轮都让人以为自己已经完整了。这正是本车道反复定价高于「缺一个检查」的失效模式。

④ 创业阶段不扩散需求. 拆成并行 job 会增加 runner 用量与维护面(120 步如何分组、谁维护分组、合并队列的必需检查名单要跟着改)。⚠️ 动的是维护者的预算,并且改的是合并队列的 required contexts —— 这两条都在人工地板上。

⭐ 四维不同向:②③ 指向改,①④ 指向先测分布再说。且无论走哪条,都触及「门禁削弱」与「合并队列必需检查」两个人工地板项。

Options

  • A. continue-on-error + 汇总步骤 —— 让全部 120 步都跑完,末尾一步汇总并在有任一失败时退出非零。⚠️ 中途每一步都显示绿,只有汇总步是真判据 —— 这本身就是一个「聚合读数」,而本仓的既有纪律是放行认门禁 job 的结论、不认聚合。要落地必须先解决这个矛盾。
  • B. 拆成 N 个并行 job —— 每个 job 内部仍 fail-fast,但彼此独立,一次运行能报出 N 条独立的红。成本:runner 用量、分组维护、以及合并队列 required contexts 名单要同步改。
  • C. 什么都不改,先测失败位置分布 —— 用历史运行统计「失败步在 120 中的位次」,得出未测量尾巴的真实期望值,再决定 A/B 是否值得。最便宜,且是唯一能把①的问号变成数字的选项。
  • D. 维持现状 —— 记录代价,接受它。10 步的尾巴不足以支付 A 或 B 的复杂度。

⛔ 本座位不背书任何一个。dev 的建议是把 ② 单独立卡带上今天的数字,已照办。

Not this card

⛔ 不是 #12211(合并队列快照状态);分诊评论已裁定两者分开。⛔ 不是 #12173 的 ①,那半已由 #11934 修复。

Refs: #12173(来源卡,其 ① 已死)· #11934 / a187fe612b(修掉 ① 的提交)· #12211(相邻但不同)

Activity

  1. os-steve commented on Aug 26, 2026

    @os-steve
    Collaborator

    Maintainer ruling recorded — Option C adopted: measure the failure-position distribution first

    Provenance: maintainer, 2026-08-26, live PM chat (decision-inbox batch 4, skills seat session session_01JANH3y7qe3MD8aLaLXci8N), verbatim: 「决策箱第 4 批 同意」 — accepting this card's presented batch-4 recommendation, which was C.

    What this rules:

    1. No change to the Lint & Repo Gates job now — no continue-on-error, no job split. Both A and B touch manual-floor territory (gate-weakening shape / merge-queue required-contexts churn / runner budget) and neither is paid for until the cost is a number instead of a question mark.
    2. The work is a measurement: over historical runs of the job, derive the distribution of the first-failing step's position among the ~120 sequential steps, and from it the expected size of the unmeasured tail per red run (the card's own 46 → 10 correction showed the single-incident estimate can be off by 4×). Deliverable = the numbers on this card, plus the re-derived step count at measurement time.
    3. After the numbers land, the card returns to the decision box for the A / B / D choice with ① priced — the measurement itself changes no CI contract and needs no maintainer floor.

    State: needs-user-decision → pm:queue in the same stroke (domain:devx lane dispatches; the measurement reads workflow run history only — no gate edits are in scope for that dispatch).


    Generated by Claude Code

  2. claude commented on Aug 31, 2026

    @claude
    Contributor

    Claim + Dispatch — R34 · ⚠️ measurement-only, no code change is in scope

    domain:devx PM seat (seat post #6023), session session_01Pk26oZ12t5N1hwGW1m1MgC.
    Branch: claude/issue-12413-gate-failure-position-measurement — create it only if you end up with files to commit. The primary deliverable of this dispatch is numbers posted on this card, not a PR.


    Zone 1 — RULING (⛔ not re-litigable, quoted verbatim, not translated)

    维护者,2026-08-26,decision-inbox batch 4,verbatim「决策箱第 4 批 同意」 — Option C adopted:

    1. No change to the Lint & Repo Gates job now — no continue-on-error, no job split. Both A and B touch manual-floor territory (gate-weakening shape / merge-queue required-contexts churn / runner budget) and neither is paid for until the cost is a number instead of a question mark.
    2. The work is a measurement: over historical runs of the job, derive the distribution of the first-failing step's position among the ~120 sequential steps, and from it the expected size of the unmeasured tail per red run (the card's own 46 → 10 correction showed the single-incident estimate can be off by 4×). Deliverable = the numbers on this card, plus the re-derived step count at measurement time.
    3. After the numbers land, the card returns to the decision box for the A / B / D choice with ① priced.

    (domain:devx lane dispatches; the measurement reads workflow run history only — no gate edits are in scope for that dispatch.)

    Binding consequences:

    1. ⛔ Do not edit .github/workflows/lint.yml. Not to add continue-on-error, not to split the job, not "while I'm here". The whole point of this ruling is that the change is unpriced.
    2. The deliverable is a comment on this card carrying the distribution, the derived expected tail size, and the step count re-derived at measurement time.
    3. This card goes back to the decision box afterwards — so write the numbers for a reader who will choose among A/B/D, not for a reader who already agrees.

    Zone 2 — PM mechanism assumptions (⚠️ measure them; falsifying one is a good outcome)

    • A2.1 — I assume the first-failing step's position is recoverable from workflow run history at useful scale. It may not be: job logs age out, and this seat has measured that the job-log tail window truncates the failing assertion. If the history you can actually reach is too thin to support a distribution, say so with the number of runs you could read — "we cannot price this from history" is a real and reportable answer that changes the decision, and it is much better than a distribution built on six runs presented as if it were solid.
    • A2.2 — I assume step position is stable enough across the window that positions from older runs are comparable to today's ~120. The card already shows this moving (the specimen slid, and the tail estimate went 46 → 10). Report the step count at both ends of your window, and if positions are not comparable, say what you normalized and why.
    • A2.3 — I assume "first failing step" is unambiguous in the logs. If a run has several failures, or the abort point is not the first red, define your rule explicitly in the writeup.

    Zone 3 — suggested route (optional; your judgment governs)

    1. Enumerate historical runs of the Lint & Repo Gates job over as long a window as the API will actually give you; record how many you got and how many you could not read.
    2. For each red run, locate the first failing step and its ordinal position; derive the distribution and the expected number of steps that never ran.
    3. Re-derive the current step count (the card's 120, from lint.yml) at measurement time.
    4. Post it here as a table with the window, N, and the readings — including the ones you could not obtain.

    ⛔ Boundaries

    • ⛔ No gate edits. No lint.yml changes.
    • ⛔ Do not touch content/docs/releases/.
    • ⛔ Never git stash — refs/stash is shared across every worktree of this repo.
    • If you do produce committed artifacts (a script, a dataset), use a dedicated worktree and a draft PR; the numbers still go on this card.

    Generated by Claude Code

  3. claude commented on Aug 31, 2026

    @claude
    Contributor

    Measurement delivered — Option C: failure-position distribution for Lint & Repo Gates

    domain:devx dev, session session_01Pk26oZ12t5N1hwGW1m1MgC. No lint.yml edit, no PR — the ruling's deliverable is numbers, and they are below.

    Headline: the expected unmeasured tail is ~65 gate steps per red run, not 10. The card's 10 is reproduced exactly, but it came from two gates that sit at ~90% of the farm — the far right tail of the distribution. Of the card's two estimates, 46 was closer to right than 10.

    But the tail is overwhelmingly green. Of 216 feature branches that went red on this farm, only 8 (3.7%) ever had a second, later gate fail after the first was fixed. That 3.7%, not the 65, is what options A and B actually buy.


    1. Re-derived step count at measurement time

    At origin/main 713f83f5d8 (2026-08-31; main moved twice during this measurement):

    reading value
    .github/workflows/lint.yml total lines 5,037
    job lint (name Lint & Repo Gates) spans lines 111-3389
    declared steps in the job 130
    real continue-on-error keys in the job 0
    job-level continue-on-error absent
    steps carrying an if: condition 0
    runtime steps the API reports 136 (130 + Set up job + 4 Post ... + Complete job)

    Card said 120 at b000ab59bb; five days later it is 130. The single textual continue-on-error hit is still a comment, now at line 2674 (card said 2451).

    Because no step carries an if:, the farm is strictly sequential and unconditional, so tail = 130 - position exactly — no conditional-skip confounder.

    2. Method — and why the log-tail truncation does not apply

    I did not read a single job log. GET /repos/{owner}/{repo}/actions/runs/{run_id}/jobs returns every job with its full steps[] array carrying a per-step conclusion. On a red run the failing gate is failure and every gate after it is literally marked skipped — so the unmeasured tail is counted, not inferred.

    Verified on run 33356347629: skipped-gate count 83 equals gates (130) minus position (47), exactly.

    ⚠️ Trap for anyone re-running this: step.number runs 1..263 with gaps and is not the position. Position must be the index within the steps array after removing the 6 synthetic steps (Set up job, 4x Post ..., Complete job).

    Cross-checked against source: for run 32816840758 the API says farm 111 / position 102 / tail 9, and lint.yml at that run's own commit has 111 declared steps with ObjectQL double limit gate at position 102. A second cross-check differed by 1 — expected, because pull_request runs execute the workflow from the merge ref, not the PR head sha. The API array is ground truth; YAML-at-head-sha is only an approximation.

    3. Window, N, and what I could NOT read

    reachable run history for lint.yml 23,475 runs
    window 2026-01-19T12:32:49Z .. 2026-08-31T06:10:01Z (run_number 4 .. 27,708)
    conclusions success 16,820 · cancelled 2,837 · action_required 2,320 · failure 1,486 · null 10 · startup_failure 2
    failed runs enumerated 1,485 (matches total_count)
    jobs fetched for them 1,485 — 0 fetch errors
    analysable gate-farm failures 466

    This is the workflow's entire life, not a sample. Retention was never the limit.

    Could not read, itemised:

    class count why
    failed runs returning zero jobs 164 total_count: 0; all pull_request, 159 of them in July, mostly changeset-release/main. No gate ever ran.
    gate-farm job failure with an empty steps array 24 job-level infra failure; no step data exists
    failed runs with no gate-farm job at all 470 before 2026-06-21 lint.yml contained only TypeScript Type Check
    action_required runs 2,320 approval-gated; 60/60 sampled returned zero jobs — never ran

    Known undercount: the run-level conclusion=failure filter misses reds hidden inside cancelled runs (concurrency cancels the run after the gate job already failed). Sampling 250 August cancelled runs found 6 (2.4%) with a failed gate-farm job — extrapolating over 2,444 August cancelled runs, roughly 59 additional red gate-farm jobs are outside my 490. Their positions looked similar; I do not expect this to move the distribution.

    ⚠️ Pagination note: the filtered runs endpoint caps at 1,000 results even when total_count says 1,485. I sliced by month to get all of them. The unfiltered endpoint has no such cap.

    4. The distribution

    Current-shape window — farm of 100 or more gates, 2026-08-24..08-31, N = 144:

    statistic position of first failing gate unmeasured tail (gates that never ran)
    mean 56.4 66.3
    median 60 61
    p10 45 49
    p25 45 61
    p75 60 82
    p90 74 82
    max 127 104

    Share of reds landing in the first 10% of the farm: 0.0% · first half: 86.1% · last 10%: 7.6%.

    Position as a decile of the farm, all 466 analysable failures:

    farm decile reds
    0-10% 13
    10-20% 48
    20-30% 46
    30-40% 79
    40-50% 71
    50-60% 16
    60-70% 26
    70-80% 5
    80-90% 92
    90-100% 70

    Sensitivity — the answer is stable at 54-72 whatever you do to the window:

    estimator N tail scaled to a 130-gate farm
    (a) every analysable failure 466 58.2
    (b) Lint & Repo Gates era, 2026-08-18 onward 243 70.5
    (c) farm of 100+ gates 144 69.9
    (d) farm of 120+ gates 126 72.4
    (e) (c) minus the two broken-on-main episode days 37 64.4
    (f) (a) minus those episode days 359 54.2
    (g) one row per (day, failing step) — per distinct defect 138 53.7
    (h) (g) restricted to farm of 100+ 30 58.7

    Call it 65, plus or minus 10.

    5. Reproducing the card's own specimen

    Measured directly, every occurrence in the history:

    gate occurrences farm position tail
    ObjectQL double limit gate 4 111-112 102-103 9 each
    WHERE-matcher conformance gate 6 54-124 48-113 6, 8, 8, 8, 10, 11

    The card's 10 is exactly right for that specimen. It is also the 93rd percentile of the tail distribution — those gates sit near the end of the farm. Generalising the specimen understated the typical cost by about 6x.

    6. ⚠️ Confound the decision must see: two broken-on-main episodes

    78% of current-shape reds are two gates, and both are single-day episodes where the gate broke on main and every open PR inherited it:

    day reds dominated by
    2026-08-28 60 56 = ADR anchors + number uniqueness (2 of them on main itself)
    2026-08-30 47 43 = Docs anchors resolve to real headings

    Estimators (e)-(h) strip these out and the answer barely moves, so the headline survives. But it means the recent red rate (0.6% to 19.0% per day, ~5% typical) is driven by main-breakage episodes, not by a steady drizzle of independent PR defects — and A and B do nothing for an inherited failure, because there is only one real defect and it is not the PR author's.

    7. The number that actually prices A and B

    The 65-step tail is an upper bound on hidden defects, not a count of them. Most of those 65 gates would have passed. What matters is how often fail-fast actually costs an extra CI round trip:

    feature/PR branches that went red on this farm at least once 216
    of those, red 2 or more times 21
    consecutive red pairs on the same branch 25
    — same gate still red (author had not fixed it) 16 (64%)
    — a later gate failed after the first was fixed 8 (32%)
    — an earlier gate now red 1 (4%)
    branches where fail-fast demonstrably cost an extra round trip 8 / 216 = 3.7%

    Median gap between the two reds: ~0.75h. 5 of the 8 were adjacent gates in the same family (Engine test-double contract gate immediately followed by WHERE-matcher conformance gate, three times).

    main is excluded from this (42 reds spread over months, not a fix-repush sequence).

    8. Runner-time pricing of A and B

    From 40 successful full-farm runs, 2026-08-28..31:

    farm wall time per run median 682s (11.4 min), range 489-861s
    first 10% of gates consume 36% of farm wall time
    first 25% of gates consume 57%
    first 50% of gates consume 76%

    The farm is heavily front-loaded: Engine query-options erasure ratchet (96s, gate 9), ESLint (56s, gate 7), Slot-lookup ratchet (51s, gate 8), Comment mask agrees with a real parser (39s, gate 19).

    • Option A — extra runner time to keep going after the first red: median 166s, mean 174s, max 412s per red run. At ~12 red runs/day that is roughly 35 extra runner-minutes/day. Cheap. A's real cost is the aggregate-reading contract conflict the card names, not compute.
    • Option B — the setup prefix (gates 1-6, through Install dependencies) is median 32s, re-paid by every split job on every run, red or green: N=4 costs about +96s/run, N=8 about +224s/run. That excludes VM allocation, which step timings do not cover, so it is a floor. B is the expensive one, and it is charged on the 95% of runs that are green.

    9. Zone 2 assumptions — verdicts

    • A2.1 (position recoverable from history at useful scale) — FALSIFIED, in the good direction. Not merely recoverable: recoverable without reading any log, so the job-log tail truncation this seat measured never bites. The whole 23,475-run life of the workflow is reachable with 0 fetch errors. ⚠️ A different thinness is real though, and the decision should hear it: the farm at today's size is about one week old, so the current-shape N is 144, of which two days supply 69%.
    • A2.2 (positions comparable across the window) — FALSIFIED. The farm grew from 7 gates (2026-06-21) to 130 (today), roughly doubling every 10 days through August: 8 (Jun) to 11 (Jul 17) to 20 (Aug 1) to 44 (Aug 10) to 76 (Aug 20) to 111 (Aug 25) to 130 (Aug 31). Job renamed ESLint to Lint & Repo Gates on 2026-08-18. Raw positions from older runs are meaningless; I normalised every window by tail as a fraction of that run's own farm and rescaled to 130, which is what the sensitivity table reports.
    • A2.3 ("first failing step" unambiguous) — HELD, cleanly. In 466 of 466 runs carrying step data, exactly one gate had conclusion failure; every gate after it was skipped. There is no ambiguity to rule on. The only exceptions are the 24 jobs with an empty steps array — job-level infra failures where no gate ran — which I excluded and counted separately rather than scoring as position 1.

    10. What this settles, and what it does not

    Settled: the tail is ~65 gates, not 10; it is stable across every window and de-duplication I tried; and the cost is ~166s of runner time, not a large compute bill.

    Not settled by measurement, and left to the decision box: the realized benefit is small — 3.7% of red branches demonstrably paid a second round trip — and the largest single driver of reds in the recent window is inherited main-breakage, which neither A nor B addresses. A reader choosing D now has a defensible number to point at; a reader choosing A can point at 166s; a reader choosing B is buying the least and paying on every green run.

    One observation the choice may want, offered as measurement and not as a proposal: the two gates responsible for 78% of recent reds sit at positions ~45 and ~59 of 130, and the four slowest gates sit at positions 7, 8, 9 and 19. Position within the farm is itself a free variable that no option A/B/D currently touches.


    Raw data and scripts are in this dev session's scratchpad (not committed — the ruling's deliverable is numbers, and a one-off pricing script would add scripts/ gate surface for a decision that has not been made). Everything above is reproducible from two endpoints: GET .../actions/workflows/224934443/runs (slice by month; the filtered endpoint caps at 1,000) and GET .../actions/runs/{run_id}/jobs.

    Generated by Claude Code


    Generated by Claude Code

  4. claude commented on Aug 31, 2026

    @claude
    Contributor

    os-dev-report

    {
      "issue": 12413,
      "status": "done",
      "branch": null,
      "pr": null,
      "premise_still_valid": true,
      "summary": "Measurement-only dispatch (Option C). No lint.yml edit, no code change, no branch, no PR - the ruling's deliverable is numbers and they are posted as comment 5474709118. Re-derived at origin/main 713f83f5d8: the `lint` job (name `Lint & Repo Gates`) spans lint.yml lines 111-3389 and declares 130 steps (card said 120 five days earlier at b000ab59bb), with 0 real continue-on-error keys (the 1 textual hit is a comment, now line 2674) and 0 steps carrying an `if:`, so tail = 130 minus position exactly. Distribution over the workflow's ENTIRE life (23,475 runs, 2026-01-19 to 2026-08-31, run_number 4 to 27,708): expected unmeasured tail is ~65 gate steps per red run (mean 66.3, median 61 in the current-shape window; 54-72 across eight window/dedup estimators), NOT 10. The card's 10 is reproduced exactly for its specimen gates (ObjectQL double limit gate: tail 9 on all 4 occurrences; WHERE-matcher: 6/8/8/8/10/11), but those gates sit at ~90% of the farm - the 93rd percentile of the tail distribution - so the specimen understated typical cost by ~6x. Of the card's own two estimates, 46 was closer to right than 10. COUNTERWEIGHT the decision box needs: the tail is real but overwhelmingly green. Of 216 feature branches that went red on this farm, only 8 (3.7%) ever had a second, LATER gate fail after the first was fixed (25 consecutive red pairs: 16 same gate, 8 later gate, 1 earlier gate). And 78% of current-shape reds are two single-day broken-on-main episodes (2026-08-28 ADR anchors 56/60; 2026-08-30 Docs anchors 43/47) which A and B do not help, because the PR author owns no defect there. Runner pricing: farm wall time median 682s; Option A costs median 166s extra per red run (~35 runner-min/day at ~12 reds/day); Option B re-pays a 32s setup prefix per split job on EVERY run including the 95% that are green. Also noted for the box, as measurement not proposal: gate ORDER is a free variable no option touches - the two gates causing 78% of reds sit at positions ~45 and ~59 of 130, and the four slowest gates sit at positions 7, 8, 9 and 19 (first 10% of gates burn 36% of wall time).",
      "tests": "No build or test - measurement only, nothing compiled and no verify lock taken. Instrument validation instead, four ways: (1) METHOD NEEDS NO LOGS, which retires the seat's log-tail-truncation hazard: GET /repos/{o}/{r}/actions/runs/{run_id}/jobs returns each job's full steps[] with per-step conclusion; on a red run every gate after the failure is literally marked `skipped`, so the tail is counted, not inferred. (2) ARITHMETIC CROSS-CHECK on run 33356347629: skipped-gate count 83 == gates 130 minus position 47, exact. (3) SOURCE CROSS-CHECK: run 32816840758 API says farm 111 / position 102 / tail 9; lint.yml at that run's own commit 860c077318 has 111 declared steps with `ObjectQL double limit gate` at position 102 - exact match. A second cross-check (run 32892002186, sha 19c50eba48) differed by 1, explained and expected: pull_request runs execute the workflow from the MERGE ref, not the PR head sha, so the API array is ground truth and YAML-at-head-sha is only an approximation. (4) SPECIMEN REPRODUCTION: the card's hand-derived 10 is reproduced from the API independently (see summary). Coverage/negative results reported rather than hidden: 1,485 failed runs enumerated (matches total_count) and jobs fetched for all 1,485 with 0 fetch errors; 466 analysable gate-farm failures. Unreadable classes counted separately - 164 failed runs return ZERO jobs (total_count 0, all pull_request, 159 in July, mostly changeset-release/main); 24 gate-farm failures have an EMPTY steps array (job-level infra failure, excluded rather than scored as position 1); 470 failed runs predate the job's existence (before 2026-06-21 lint.yml held only `TypeScript Type Check`); 2,320 action_required runs never ran at all (60/60 sampled return zero jobs). KNOWN UNDERCOUNT declared: run-level conclusion=failure misses reds hidden inside cancelled runs; sampling 250 August cancelled runs found 6 (2.4%) with a failed gate-farm job, extrapolating to roughly 59 more reds outside my 490. PLATFORM TRAPS HIT AND RECORDED: (a) step.number runs 1..263 with gaps and is NOT the position - position must be the array index after removing the 6 synthetic steps (Set up job, 4x Post ..., Complete job); (b) the FILTERED runs endpoint caps at 1,000 results even when total_count says 1,485, so the failure set was collected by month-slicing (the unfiltered endpoint has no such cap); (c) I initially mis-mapped the workflow id from the listing and swept 224934441 (codeql.yml) - caught by checking the returned `path` field before analysing, and re-run against the correct id 224934443; no wrong-repo/wrong-workflow number reached the card. Channel: REST only (repo-scoped probe returned HTTP 200, core limit 15,000/hr), plus local git for lint.yml at pinned shas. Zero MCP GitHub calls. Deliverable comment 5474709118 was read back in full after posting: body byte-identical to what I wrote, all key figures present, nothing eaten by the sanitizer; the only delta is a second signature footer the platform appended itself, which I left alone.",
      "mcp_calls": "0 - REST plus local git throughout; no MCP GitHub call was made",
      "open_questions": [
        {
          "question": "The measurement surfaced a lever none of A/B/D contains: gate ORDER within the farm. The two gates causing 78% of recent reds sit mid-farm (positions ~45 and ~59 of 130), which is precisely why the measured tail is ~65; and the farm is heavily front-loaded for time (first 10% of gates burn 36% of wall clock). Should the decision box consider ordering as an option alongside A/B/D, or is it out of frame for this card?",
          "options": [
            "A. Keep the box to A/B/D as ruled; record ordering as a separate observation for a future card.",
            "B. Add an ordering option to this card's box (cheap-and-broad gates first, or move the two dominant gates earlier so their tail is small).",
            "C. Treat ordering as out of scope entirely - the farm order encodes dependencies that this measurement did not verify."
          ],
          "recommendation": "A - the ruling scoped this dispatch to producing a number and explicitly returns the card to the box for A/B/D. Ordering is a real finding but it is a NEW option with its own unpriced cost (I did not verify whether the current order encodes step dependencies), and folding it in now would repeat exactly the mistake the ruling corrected: choosing before the cost is a number. Recording it as an observation preserves it without widening an unpriced decision."
        }
      ],
      "out_of_scope_findings": []
    }

    Generated by Claude Code


    Generated by Claude Code

  5. claude commented on Aug 31, 2026

    @claude
    Contributor

    ⚖️ 测量完成 —— 回箱,①已定价。卡的头条数字错了 6 倍,而结论仍落在 D,但理由必须换掉。

    domain:devx PM 席(座位贴 #6023),session session_01Pk26oZ12t5N1hwGW1m1MgC。裁决要求的交付物是数字,已由 dev 贴在上方(评论 5474709118)。⛔ 未改 lint.yml,⛔ 无 PR,⛔ 无分支 —— 测量型派发按裁决只产出数字。

    os-decision-facets

    先说本席独立核过的三条(⛔ 不是转述 dev 报告)

    读数 核验
    lint job 现有 130 步(卡记 120) 本席自己数:.github/workflows/lint.yml 第 111–3389 行,- name: 130 条
    真 continue-on-error 0 个 唯一文本命中在 2674 行,是注释
    ⇒ 尾巴 = 130 − 位次,精确 无 if: 步,无 continue-on-error,不存在部分执行

    ①(卡上唯一的问号)现在是数字

    每次红,平均有 ~65 道门禁从未跑过(均值 66.3 / 中位 61;八种窗口与去重估计量落在 54–72)。样本:23,475 次运行,2026-01-19 → 2026-08-31。

    ⭐ 卡自己的 10 被精确复现了,但它是个离群值 —— ObjectQL double limit 门禁四次出现尾巴都是 9,WHERE-matcher 是 6/8/8/8/10/11;这两道门恰好坐在农场 ~90% 处,即尾巴分布的第 93 百分位。⇒ 卡用一个第 93 百分位的样本去代表典型值,低估了约 6 倍。卡自己的两个估计里,被它划掉的 46 反而更接近真相。

    ⇒ D 的正文理由「10 步的尾巴不足以支付 A 或 B 的复杂度」已被证伪。 若仍选 D,必须换理由,否则这张卡将来会被引用为一个假前提。

    ⚠️ 但同一份测量给出一个强反向配重 —— 而它恰好救回了 D 的结论

    • 尾巴大,却几乎不咬人:216 条在本农场红过的特性分支里,只有 8 条(3.7%) 在修掉第一道红之后、又被一道更靠后的门禁拦住。25 对连续红中 16 对是同一道门,8 对更靠后,1 对更靠前。
    • 78% 的当期红根本不是 PR 的锅:两次单日 main 破损事件(2026-08-28 ADR anchors 56/60;2026-08-30 Docs anchors 43/47)。A 和 B 对这类一点忙都帮不上 —— 作者名下没有缺陷。

    ⇒ 「未测量的尾巴」很大(~65),但「因此多付一个来回」很罕见(3.7%)。 这两个数字必须一起读;单看任何一个都会选错。

    定价(runner)

    成本
    农场墙钟中位 682s
    A 每次红多付中位 166s(约 12 次红/天 ⇒ ~35 runner-分钟/天)
    B 每个拆出的 job 每次运行重付 32s 启动前缀 —— 包括 95% 的绿跑

    ⚠️ A 另有一个与成本无关的硬阻碍,卡自己已点出:A 的判据是末尾汇总步,而本仓纪律是放行认门禁 job 的结论、⛔ 不认聚合读数。要落地 A 必须先解决这个矛盾 —— 那是纪律问题,不是钱的问题。

    四维

    ① 实际业务需求:真实但小。3.7% 的分支会因尾巴多付一个来回;78% 的红与此无关。
    ② 平台长远合理性:130 步单 job 顺序跑本身是脆的,但 A 用聚合读数换可见性,与既有纪律直接冲突;B 把成本摊到 95% 的绿跑上,为 5% 服务。
    ③ 防 AI 写代码犯错:尾巴让 agent 修完一道红以为完事 —— 但 3.7% 说明这个错觉极少变成真实返工。A 会引入一个更危险的错觉:每一步都显示绿。
    ④ 创业阶段不扩散需求:A 和 B 都要动 CI 契约与合并队列 required contexts;为 3.7% 付这个,不划算。

    推荐 D(维持现状),但正文理由整段替换为:"尾巴经实测约 65 步,远大于立卡时的估计;然而红后被更靠后门禁二次拦截的分支只有 3.7%,且 78% 的红源于 main 破损而非 PR 缺陷 —— 因此尾巴虽大,代价罕见,不足以支付 A(与聚合读数纪律冲突)或 B(向 95% 的绿跑收费)。"

    置信缺口(逐条,⛔ 不藏)

    • 已声明的漏计:run 级 conclusion=failure 看不见藏在 cancelled 里的红。抽 250 次八月 cancelled 运行,6 次(2.4%) 含失败的农场 job ⇒ 外推约多 59 次红落在这 490 之外。
    • 不可读类别单列:164 次失败运行返回零 job;24 次农场失败 steps 为空(job 级基础设施失败,已排除而非记为位次 1);470 次早于该 job 存在;2,320 次 action_required 从未运行。
    • 本席未独立复算分布本身 —— 我核的是 130 步、0 个 continue-on-error、以及方法是否成立;66.3 这个均值我采信 dev 的枚举。

    ⭐ 方法值得单独记一笔,它退役了本席警告过的一个风险:我在派发令里担心「日志尾窗会截断失败断言」。dev 根本没读日志 —— GET /actions/runs/{id}/jobs 直接给出每个 job 的 steps[] 与逐步 conclusion,红跑上失败之后的每一步字面标着 skipped,尾巴是数出来的,不是推出来的。并做了算术自洽(83 == 130 − 47)与对源交叉验证。⇒ 我的 A2.1 假设被以更好的方式绕开,而非勉强满足。

    一个新杠杆 —— ⛔ 本轮不并入本箱

    测量顺带发现:门禁顺序是所有选项都没碰的自由变量。造成 78% 红的那两道门坐在 130 步的第 ~45 与 ~59 位(这正是尾巴 ~65 的来源);而最慢的四道门坐在第 7、8、9、19 位 —— 前 10% 的门烧掉 36% 的墙钟。

    ⛔ 不并入本卡的箱,与 dev 的建议一致:它是一个新选项,且自带未定价的成本(当前顺序是否编码了步骤依赖,本次未验证)。把它现在折进来,恰好会重犯这张裁决当初纠正的那个错误 —— 在成本还是问号时就先选。另立观察卡记录。


    Generated by Claude Code

  6. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    ⚖️ 裁决记录 — D:维持现状;理由整段替换为实测版

    出处:维护者 2026-08-31,live PM chat,总监席第 7 场决裁批 #15,逐字:「同意」(对本批按呈报推荐整体放行;本卡呈报推荐为 D)。录裁:director seat, session session_01KGtaLpkW1mycWgkbSb3H6t。接续 2026-08-26「决策箱第 4 批 同意」的 C 裁定 —— 测量已交付(评论 5474709118),本裁是回箱后的 A/B/D 终选。

    裁定内容

    D —— Lint & Repo Gates 维持首红即停的现状。 本卡正文的 D 理由(「10 步的尾巴不足以支付复杂度」)已被测量证伪,⛔ 不得再被引用;替换理由如下(即裁决的正式依据):

    尾巴经实测约 65 步,远大于立卡时的估计;然而红后被更靠后门禁二次拦截的分支只有 3.7%,且 78% 的当期红源于 main 破损事件而非 PR 缺陷 —— 因此尾巴虽大,代价罕见,不足以支付 A(每红 166s 且与「放行认门禁 job 结论、⛔ 不认聚合读数」的既有纪律正面冲突)或 B(向 95% 的绿跑每次收 32s×N 启动费,并牵动合并队列 required contexts)。

    • A、B 均否决,依据上段;两者的人工地板顾虑(门禁削弱形状 / 合并队列契约 / runner 预算)与实测定价同向。
    • 门禁排序杠杆(最慢四道门在第 7/8/9/19 位烧 36% 墙钟;两道主红门在第 ~45/~59 位)按 dev 建议 A 处理:⛔ 不并入本箱 —— 它是自带未定价成本的新选项(顺序是否编码依赖未验证)。另立 finding 观察卡记录,编号随后回贴。

    状态转移(同笔)

    决定已做、无剩余工作 ⇒ 关卡(completed),needs-user-decision 随关摘除(tooling / domain:devx 为归属标签,留)。测量数据与替换理由以本卡评论为正典存档。


    Generated by Claude Code

  7. huangyiirene commented on Aug 31, 2026

    @huangyiirene
    Collaborator

    Follow-up to the ruling above: the gate-ordering observation is already on file as #13690 (filed by the domain:devx lane before this ruling landed) — dedup hit, no second card created. That card is the recorded home for the ordering lever.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions