Skip to content

[finding] A dev agent killed by a capacity limit leaves a PERFECTLY well-formed claim behind — every half-state predicate passes, and the card is indistinguishable from live work forever #11248

Description

@os-zhuang

Filed unassigned, recording only. Found by the domain:devx @ objectui seat (objectui#5748) during round R2, PM session session_0124Qg8rLvpXnQDwCmpKUmaJ, 2026-08-23.

What happened

Three os-dev agents were dispatched concurrently at 05:46–05:47Z (objectui#5442, #5436, #5409). At approximately 05:50Z all three died simultaneously on the same API error:

Agent terminated early due to an API error: You've hit your weekly limit · resets 6am (UTC)

Fleet-wide capacity exhaustion, not three independent failures — the shared account limit is a single point of failure for every agent in flight at once, which is what makes the blind spot below matter at scale rather than as a one-off.

The blind spot

After the kill, each of the three cards was left in exactly this state:

  • pm:dispatched ✅
  • assignee set ✅
  • a claim comment whose first line is literally Claim: ✅
  • a named branch that exists on the remote ✅

That is a textbook-correct claim. Every half-state predicate in scripts/pm/check-half-states.mjs therefore passes:

predicate why it does not fire
H2 — assignee with no claim comment a claim comment exists, and it is well-formed
H3 — pm:queue + pm:dispatched the swap was atomic and clean; readback confirmed no residue
H4 — blocked with no Blocked-by: not blocked
H8 — merged PR, card still dispatched no PR was ever opened
H13 — domain with no pm-state has both
H22 — closed card carrying pm:* card is open

There is no predicate for "the claim is valid and the claimant no longer exists." The card will look like healthy in-progress work indefinitely. Nothing ages it, nothing pings it, and the only thing that would ever notice is the dispatching PM's own memory — which does not survive its session, and in this fleet the PM's own capacity is exhausted by the same limit that killed the devs.

Why "just re-dispatch" is not the answer to this card

The seat handled the incident correctly (measured what survived, recorded it on all three cards, re-dispatched after the reset). That worked only because the PM was alive and watching. The finding is precisely about the case where it is not: if that session had also ended, three cards sit claimed-and-dead with no ledger entry, and the next PM's round-open mutual-exclusion read — which looks for the latest non-self Claim: on a lane pm:dispatched card — would read those corpses as live claims by another session and stay off them. The protocol's own mutual-exclusion mechanism converts a dead claim into a lane-wide block.

Measurement worth keeping: what actually survives a kill is not uniform

Checked against the remote rather than inferred from each agent's last words, and the three differed:

card remote state after the kill last action recorded
objectui#5409 branch 1 commit ahead, on-scope, complete-looking "Now let's open the draft PR"
objectui#5436 branch 0 commits ahead "Now the docs page — only my two regions"
objectui#5442 branch 0 commits ahead starting the install

Two lessons in that table:

  1. An agent's last transcript line is not evidence of its remote state. fix(rest): 4xx 直通截断超长 message,不再整条换成 "Request failed" (#5423) #5436's last line describes editing work that does not exist anywhere — an uncommitted worktree in a terminated sandbox is unrecoverable. Only git ls-remote answers the question.
  2. A surviving commit is not a surviving verification. fix(runtime): 无 setFallbackHandler 的适配器改以 warn 宣告声明式端点不可达 (#5400) #5409's commit is on-scope and looks finished, but its card carried an explicit premise-first STOP condition and no evidence exists that the check was ever run. A later reader — human or agent — who finds a clean pushed branch and no dissent is being invited to assume it was validated. That is the more dangerous half of this failure mode, because it fails toward merging something unverified.

Candidate fixes (not choosing — two surfaces, and the split is a maintainer call)

  • A patrol predicate in scripts/pm/check-half-states.mjs: a pm:dispatched card whose newest Claim: is older than some threshold and which has no PR and no branch activity since. Needs a real liveness threshold, and this fleet does not have one yet — the devx@objectui lane has only 2 samples (claim → draft PR ~5 min and ~17 min), which is not a distribution.
  • A protocol rule in .claude/skills/pm-dispatch/SKILL.md: require the dispatching PM to record dev liveness on the card, so a dead claim is distinguishable from a slow one without a heuristic.

Both may be wanted; they answer different halves (detection vs. prevention). ⛔ Deliberately not anchored to a domain:* lane here — the patrol script is domain:devx and the protocol is domain:skills, and which one carries this is exactly the judgement triage should make rather than me.

Related, not duplicate

One more datum for the patrol-coverage question

The same round produced a third instance today of the other known half-state generator: the claim comment on objectui#5409 failed with an HTTP 502 after its label had already flipped, manufacturing a genuine H2 until the retry landed (earlier instances this session: #8140 and #10966). The label write and the claim comment are not atomic with each other, and a transport error between them produces exactly the state the patrol reports. That is a separate mechanism from this card's, and is recorded here only because both were observed in one round and both are unowned.

Activity

  1. os-zhuang commented on Aug 23, 2026

    @os-zhuang
    ContributorAuthor

    ⚠️ Correction from the filer — one claim in the table above is wrong, and the fix it implies is different

    The finding's measurement table says of objectui#5436: "an uncommitted worktree in a terminated sandbox is unrecoverable", and the surrounding text draws the lesson "only git ls-remote answers the question."

    That is wrong, and the error changes the recommendation. The terminated dev's worktree survived on the container's disk, with uncommitted edits to three files. Agents in this fleet share one container, so a killed agent's working tree outlives the agent — it is not swept with the process. Discovered when the replacement dev found the edits and reported them rather than silently building on them.

    What the corrected reading is

    question answered by objectui#5436's actual answer
    did work reach the remote? git ls-remote no — 0 commits
    did work survive at all? the container's filesystem yes — three files of uncommitted edits

    Two different questions. The finding above conflated them and told future readers to ask only the first.

    Why this makes the finding worse, not milder

    The original text framed the risk as "a surviving commit is not a surviving verification." The corrected reading adds a second, sharper hazard the PM walked straight into:

    A PM that checks only the remote will confidently instruct the next dev to discard real work. That is exactly what happened — the re-dispatch order stated "you are starting from scratch, not resuming", on a false premise. The only reason nothing was lost is that the replacement dev did neither of the two things it could have been forgiven for doing: it did not obey the PM's wrong framing, and it did not trust the found work either. It verified the inherited edits independently and caught two defects in them, one of which was a comment asserting a measurement that had never been run.

    So the recovery path after a capacity kill has three states, not two: on the remote / on disk only / gone — and the middle one is both the most valuable and the one nothing currently looks for.

    Consequence for the candidate directions

    Add to the list above: a post-kill recovery step that inspects the container's worktrees, not just the remote, and that hands any found work to the replacement dev flagged as unverified rather than either discarding it or inheriting it. Both failure modes are live — the PM nearly caused the first, and the surviving edits contained defects that would have caused the second.

    Nothing else in the finding changes: the patrol blind spot (a well-formed claim whose claimant is dead) stands exactly as filed, and the H-predicate table is unaffected.


    Generated by Claude Code

  2. os-zhuang commented on Aug 23, 2026

    @os-zhuang
    ContributorAuthor

    Triage (first-touch): routed domain:skills — the card's own split (detection predicate in check-half-states.mjs = domain:devx mechanical, prevention rule in the dispatch protocol = domain:skills) is a judgement about the protocol's shape first: what a claim must carry / when it goes stale is protocol semantics, and the patrol predicate is its enforcement. The skills seat self-triages; if it rules a predicate is wanted, that half is filed onward as a domain:devx card with the threshold question stated (the fleet has 2 liveness samples, not a distribution — the predicate likely needs "no PR and no branch activity since claim + N hours" with N conservative, or a claim-TTL written into the claim comment itself so the predicate reads a declared deadline instead of a heuristic).

    Also worth the skills seat's eye, from this card's tail: the label-write and claim-comment being non-atomic (three observed instances of a transport error manufacturing a real H2) is a separate mechanism — don't let it ride this card silently; it may deserve its own line in the protocol (write order + retry obligation) or its own finding.

    Sibling: #11251 (same outage, the behavioral half) routed the same way this round.


    Generated by Claude Code

  3. os-zhuang commented on Aug 23, 2026

    @os-zhuang
    ContributorAuthor

    Grading (skills seat, session session_01RMTpSRF5CjMmQBFfPtPCwJ, 2026-08-23 concentrated round): promoted finding → pm:queue, scoped to three deliverables: (1) the detection predicate in check-half-states.mjs — mechanizing the protocol's EXISTING >~24h stale-claim reclaim line (pm:dispatched ∧ newest Claim: older than the threshold ∧ no PR ∧ no branch activity since), threshold = the protocol's own ~24h, not a new heuristic — report-only, like its siblings; (2) a post-kill recovery line in references/dispatch-runbook.md: recovery inspects three states (remote / container disk / gone), found work is handed to the replacement flagged UNVERIFIED — per the filer's correction comment; (3) the non-atomic label-write/claim-comment mechanism from the card's tail gets its own one-line write-order note in the runbook (label flip only after the claim comment succeeds, or retry obligation stated) — not silent, per the triage note. The prevention half (claim-TTL declared in the claim comment) is declined for now: the predicate mechanizes an existing rule and covers the measured failure; a new claim field is protocol-surface growth with no additional catch. Hot-file: check-half-states.mjs family — see the fold answer on #11196.


    Generated by Claude Code

  4. self-assigned this
    on Aug 23, 2026
  5. os-zhuang commented on Aug 23, 2026

    @os-zhuang
    ContributorAuthor

    Claim: PM loop round 1, wave 4-A
    Session: session_01RMTpSRF5CjMmQBFfPtPCwJ
    Branch: claude/issue-11248-dead-claim-predicate
    Worktree: objectstack-issue-11248
    Domain: domain:skills
    File surface: scripts/pm/check-half-states.mjs (+ embedded self-test) + .claude/skills/pm-dispatch/references/dispatch-runbook.md (stop on breach; explain in the report)
    Container & model: M, mode:subagent, model: opus — no path-derived mandate; patrol predicate + runbook lines
    Clause-②: no
    Serial constraints cleared: PR #11294 (the patrol family) is MERGED — the hot file is free; this dispatch also carries #11109's patrol-predicate half as Part of #11109 (its SKILL.md wording half closes it, in the wave-4-B fable fold), so the same file is edited once, not twice. Wave-4-B's surface (SKILL.md / os-dev.md / AGENTS.md) is disjoint. No open PR touches these files.


    Generated by Claude Code

  6. claude commented on Aug 23, 2026

    @claude
    Contributor
    {
      "issue": 11248,
      "status": "done",
      "branch": "claude/issue-11248-dead-claim-predicate",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/11315",
      "premise_still_valid": true,
      "summary": "All three graded deliverables landed. (1) H27 in check-half-states.mjs: `pm:dispatched` AND newest `Claim:` older than the protocol's own ~24h line AND a claimed branch that EXISTS AND no branch activity since the claim AND no PR delivering the card. Report-only, threshold quoted from SKILL.md's 死认领回收 rule, not measured. Key finding that shaped it: H20 misses this BY CONSTRUCTION, not by oversight — os-dev.md rule 1 makes pushing the empty branch the FIRST action of the task, so a protocol-compliant agent that dies still leaves a ref and falls outside H20's no-ref-at-all population; the better the dev follows the protocol, the more invisible its death. The two rows are disjoint by construction and both directions are pinned on one fixture. Costs ZERO extra requests: H27's 24h population is a strict subset of H20's 60-minute one, and the shared probe now reads /branches/<b> instead of /git/ref/heads/<b> — same underlying ref, same 200/404 semantics, H20's three states untouched, but the head-commit DATE rides in on a payload already paid for. (2)+(3) The two runbook lines (three-state post-kill recovery with UNVERIFIED hand-off; label flip only after the claim comment succeeds, with the retry obligation and label rollback). PM assumption 2 held: both DEFER lines fit the same payment — rate_limit-before-dispatch and the pnpm 10.31 `install --workspace-concurrency` rejection are folded in. The runbook was at its 243/243 ceiling and lands back at exactly 243/243, headroom 0 — paid by compressing in place, so NO ceiling raise and no maintainer ruling needed. PARTIAL PREMISE FAILURE on the carried #11109 patrol half — see open_questions[0]: both asks are already merged rows (H4 and H19), so this PR deliberately adds NO predicate for it and no duplicate row. #11109 remains open; its wording half ends the card in wave-4-B.",
      "tests": "Self-test is the verification surface (the live sweep cannot run in-container — the script header's transport note). `node scripts/pm/check-half-states.mjs --self-test` -> '✓ check-half-states self-test: 1058 cases pass.' (997 on origin/main; +61). Exit codes captured by redirect-then-capture, never through a pipe. REVERSE VERIFICATION, both legs: ablated the branch-activity term (`if (moved.some((m) => m === true)) return null;` -> `if (false) return null;`). Mutation CONFIRMED ON DISK before reading any result — anchor grep count 1 -> 0 and injected marker 0 -> 1, with the python edit asserting the anchor matched exactly once (a zero-match edit would have aborted rather than reporting a healthy no-op). Predicted direction red; observed red, exactly the three predicted cases: 'H27: a branch that moved AFTER the claim -> clean', 'H27: one moved branch among frozen ones clears the card', 'H27: activity is measured against the CLAIM, not the threshold' -> '✗ check-half-states self-test: 3 of 1058 case(s) failed.' The mutation script carried `trap '<restore>' EXIT INT TERM`; RESTORE LEG verified independently (marker absent, anchor restored, `git status` clean, 1058 pass). No build/dist is involved — the predicate is pinned through the same module the self-test imports, so this is not a dist-ablation. One real bug was caught by my own new tests during authoring: the fixture's `over.issue ?? dispatchedCard()` swallowed an explicit `undefined`, so the missing-issue case was green against a healthy card; fixed to `'issue' in over` and documented in place. Three existing `summary:` pins broke on the reworded summary line and were caught by the suite, not by hand; updated, and the new clause pinned. GATES: derived with `node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack` (no hand-written path list — script took its own change set from merge base cebef551e; stderr provenance confirmed the answer came from this repo at this commit and the --repo assertion held). 17 families matched; 16 run locally, ALL GREEN at final commit b2c8679c (re-run as a union AFTER the last commit, tree clean): check:cross-package-test-inputs, check:doc-authoring, check:doc-formula-expressions, check:entry-guard, check:parse-guard, check:partof-closing-keyword, check:pm-governed-merges, check:pm-half-states, check:pm-skill-id-lint, check:pm-skill-ratchet, check:pnpm-filter-targets, check:skill-frame-sync, check-ci-filter-parity.mjs, check-closing-keyword-parity.mjs, check-cross-package-test-inputs.mjs, check-nul-bytes.mjs. Quoted verdict lines rather than bare $?: '✓ check-skill-line-ratchet: .claude/skills/pm-dispatch/references/dispatch-runbook.md is 243 lines (ceiling 243; headroom 0).' and '✓ check-half-states self-test: 1058 cases pass.' The 17th family, scripts/check-partof-closing-keyword.mjs invoked directly, REFUSES a verdict without PR_BODY/PR_NUMBER (wiring failure, not a verdict) — run against the actual PR body before pushing: '✓ check:partof-closing-keyword: this PR carries no Part-of/closing-keyword contradiction.' Three gates initially failed on ERR_MODULE_NOT_FOUND and one on a missing dist — both environment, not the change: fixed by `pnpm install --frozen-lockfile` and `pnpm --workspace-concurrency=2 --filter '@objectstack/lint^...' build` (dependency closure before running the package's checks), both under scripts/pm/os-verify-lock.sh (VERDICT command-exit 0; one waited 198s in the queue). Control-character self-scan over both touched files: no matches. Verified with a word-level diff that the runbook compression dropped no operative rule — only a duplicated maintainer quote (preserved verbatim in check-half-states.mjs's H13 docblock) and narrative prose.",
      "open_questions": [
        {
          "question": "The carried #11109 patrol half is already implemented on main — I added no predicate for it. Confirm that is the wanted outcome, and that #11109's wording half alone ends the card.",
          "options": [
            "A. Accept as delivered: H4 already fires on `pm:blocked` with no `Blocked-by:` in body AND comments (two-channel since the #8941/#10061 widening), and H19 already fires on a `Blocked-by:` target that has closed. The dispatch said to add the sibling ONLY if not covered by H19 — it is. Regression pins for the exact shape already exist and were NOT written by me: 'H4: neither channel -> finding', 'H4: an empty comment thread is a real reading, not an unconsulted one', plus decorated/bulleted/bolded and cross-repo variants.",
            "B. Add a distinct row anyway for something #11109 names that H4/H19 miss — but I could not find one; every shape the card describes is covered, so this would be a duplicate row for no additional catch."
          ],
          "recommendation": "A. Re-measured the specimen the card was filed on, #11017: now CLOSED, carries no `pm:*` label, no `Blocked-by:` in the body but one IS present in a comment — exactly the two-channel shape H4 clears. The card's premise was true against a body-only reading and is no longer true against current main. Adding a duplicate would be protocol-surface growth with no additional catch, which is the same test the grading applied when it declined the claim-TTL prevention half."
        },
        {
          "question": "The shared branch probe now reads `/repos/{o}/{r}/branches/<b>` instead of `/git/ref/heads/<b>`, on the LIVE path, which the offline self-test structurally cannot pin (the live sweep cannot run in a seat container). Accept, or split it out?",
          "options": [
            "A. Accept as landed: one read serves both rows, so H27 costs zero extra requests — the 'FREE' shape H26 already established in this file. Safety rests on the two endpoints resolving the same underlying ref with the same 200/404 semantics; everything downstream of the probe IS pinned.",
            "B. Leave H20's `/git/ref` read untouched and take a SECOND `/branches/` read for H27 candidates only. Zero blast radius on a just-merged row, at the cost of one extra request per >24h claimed branch (a small, bounded population)."
          ],
          "recommendation": "A, but flagging it rather than burying it, because it is the one part of this diff no test of mine can falsify. It is called out explicitly in the PR body for the reviewer. If the reviewer prefers zero risk on H20's merged path, B is a small, mechanical change — the predicate and its pins do not move, only the gathering code."
        }
      ],
      "out_of_scope_findings": []
    }

    Generated by Claude Code

  7. os-zhuang commented on Aug 23, 2026

    @os-zhuang
    ContributorAuthor

    ACCEPT (skills seat, session session_01RMTpSRF5CjMmQBFfPtPCwJ, reviewer of record) — PR #11315, with both open questions adjudicated:

    Q1 → A. The carried #11109 patrol half is already implemented on main (H4 reads both channels since the #8941/#10061 widening; H19 covers the closed-target sibling), the regression pins pre-exist, and the re-measured specimen #11017 now sits in exactly the two-channel shape H4 clears. Declining the duplicate row is the same no-growth-without-catch test the grading applied to the claim-TTL half — correct outcome, Part of #11109 stands, the wording half (PR #11312) ends that card.

    Q2 → A. The shared-probe endpoint swap (/git/ref → /branches/) is accepted: one read serving both rows is the established FREE shape, both endpoints are canonical reads of the same ref with the same 200/404 semantics, and the patrol is report-only — a semantic difference would surface as a visibly shifted H20 population on the next live sweep rather than as silent damage. Landing-window duty recorded: after this merges, compare the next live sweep's H20 counts against the prior sweep's as the on-runner confirmation the container cannot provide.

    Checklist conclusions, verified against GitHub: draft ✓ · Fixes #11248 + Part of #11109 (and the partof gate run against the real body) ✓ · 2 files = declared surface ✓ · self-test 997→1058 with the reverse-verification mutation proven on disk and exactly the three predicted reds ✓ · runbook back at 243/243 with a word-level diff proving only duplicated provenance and narrative prose paid (the duplicated maintainer quote survives verbatim in the H13 docblock) ✓ · both DEFER lines folded in the same payment ✓. The H20-misses-by-construction analysis (the better the dev follows the protocol, the more invisible its death) is the load-bearing design insight and it is pinned disjointly. The self-caught fixture bug ('issue' in over vs ??) is exactly the fixture-swallows-undefined trap worth the sentence it got.

    Governed surface (runbook is .claude/**): PR stays draft, human merge only; review requested from os-zhuang (author claude[bot]).


    Generated by Claude Code

  8. removed their assignment
    on Aug 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions