Skip to content
This repository was archived by the owner on Oct 1, 2026. It is now read-only.

Delete elapsed-time termination of agent turns and jobs (#88) - #90

Merged
lukemaj merged 3 commits into
mainfrom
fix/88-no-elapsed-worker-kill
Sep 24, 2026
Merged

lukemaj merged 3 commits into
mainfrom
fix/88-no-elapsed-worker-kill

Conversation

@lukemaj

@lukemaj lukemaj commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Implements #88.

Outcome: productive agents stay running regardless of total elapsed time. Elapsed-time killing is deleted, not increased: no per-turn or per-job deadline exists anywhere in the runner.

What changed (base 77d0010):

  • runner/harnesses.py: timeout_for() returns None for every kind; per-harness timeout tables removed; effective_stall_window_secs() passes the policy window through unclamped (legacy slot ignored).
  • runner/core.py: _durable_run records the legacy timeout value but never enforces it (new rows store NULL); controller adopts the supervisor result with no wait deadline; wait helpers adopt live work until ownership/cancel/terminal; recover_one no longer cancels or drains jobs by age (legacy positive timeout_secs ignored).
  • runner/supervisor.py: no turn deadline on the generic CLI path or the owned-server drive; stall window unclamped; stall probe on the fixed policy budget; serve startup/health on a fixed 120s readiness budget; explicit cancel (rc143) and real failures unchanged.
  • Policy 2.7.0 -> 2.7.1 with Remove elapsed-time termination of active agents and jobs #88 provenance; skill reference regenerated from policy (not hand-edited); GLOSSARY Silence window no longer asserts an outer budget; RUNNER.md and CLI --timeout-secs document the legacy compatibility treatment (accepted, recorded, never enforced).
  • Tests: new tests/test_issue88.py (deterministic tests, seconds not minutes); updated tests that encoded the rejected deadline (test_issue73 bound/boundary, test_faults hung-turn + job-timeout, test_grok hung worker, test_stall33 header). Cancellation, ownership, stall, failure, and legacy-drain regressions preserved.

Correction da15150 (independent review of 77d0010, REQUEST_CHANGES, 2 findings):

  • runner/supervisor.py: the OpenCode drive loop could poll forever when the session went idle after activity without a terminal assistant result. The non-busy branch now reuses the existing activity tracker and stall/probe/abort machinery (prompt-waiting state exempt, live part/tool activity keeps the turn alive), so genuine silence past the policy window ends the turn as stalled with the existing signal/evidence shape. No elapsed deadline restored.
  • runner/controller.py: Grok supervisor loss (rc 125, CLI child alive but unsupervised) now blocks as ownership like 124/143 instead of falling through as an ordinary worker failure, so no new attempt queues behind the unsupervised child. Timeout language removed from the comment; OpenCode path and all other rc behavior unchanged.
  • tests/fakes.py: faithful idle_incomplete (streams incomplete messages while busy, then idle) and idle_empty (busy, then idle with no message) fake modes.
  • tests/test_issue88.py: idle-incomplete/idle-empty stall tests (stalled_retry, stream_silence evidence, idle_confirmed, stalled reports); Grok rc125 classification tests.

Correction 3a24647 (coordinator proof-gap review of da15150):

  • tests/test_issue88.py only: faithful supervisor-loss drill. A real core._durable_run starts the fake Grok hold child in a thread; after the row records both identities and the worker logs its spawn, only the exact recorded supervisor pid is SIGKILLed. The durable run returns rc 125 with the owned child live and ownership live; that authentic result fed through controller.run_implementation blocks as ownership (single controller call, exactly one worker spawn ever, no escalation). Targeted owned-child cleanup by exact pid, explicit cancel works, unrelated sleep process survives. Mutation-checked: with the controller fix reverted the drill fails (implementation_failed instead of blocked). Also fixed a log-read race (spawn gate before the kill) found by the full suite under load. No production change.

Proof (exact candidate 3a24647):

  • Targeted: tests.test_issue88 (14 tests) OK; tests.test_grok + tests.test_stall33 (66 tests) OK; faithful drill stable across repeats and alongside GrokDurable/GrokControllerInjected.
  • Full: python3 scripts/test.py -> 599 tests OK (skipped=1) in ~335s, log-verified (one earlier run caught the fixed log-read race; final run green).
  • python3 -m compileall -q runner scripts tests clean.
  • Version stays 0.29.2 from 0.29.1 and policy 2.7.1 (correction before delivery, not another release; no VERSION/CHANGELOG/mirror change). Note: the pre-commit version gate was bypassed with --no-verify for these correction commits only, per the explicit no-repeat-bump instruction; message conventions kept.
  • Prior da15150/77d0010 proof does not stand in for this candidate; the proof above is on 3a24647.

No merge, release, or install. Independent exact-candidate re-review still required (coordinator-owned).

…ership

Finding 1: the OpenCode drive loop could poll forever when the session
went idle after activity without a terminal assistant result. The
non-busy branch now reuses the existing activity tracker and
stall/probe/abort machinery when the turn is not terminal and not
prompt-waiting, so genuine silence past the policy window ends the
turn as stalled. No elapsed deadline is restored; live part/tool
activity keeps the turn alive.

Finding 2: core._durable_run rc 125 (supervisor lost with the CLI
child still alive) fell through the Grok ownership check as an
ordinary worker failure, risking a new attempt queued behind an
unsupervised child. rc 125 now blocks as ownership like 124/143,
with timeout language removed from the comment.

Tests: faithful idle_incomplete/idle_empty fake modes plus
stalled_retry assertions; Grok rc125 ownership/blocked assertions
with no-duplicate-writer, unknown-process, cancel, and rc1
preservation checks.
Replace the log-read race with a spawn gate and add the reviewed
trigger faithfully: a real core._durable_run loses its supervisor
(SIGKILL on the exact recorded supervisor pid) while the owned
fake-Grok hold child stays live. The durable run returns rc 125 with
the child live and owned; feeding that authentic result through
controller.run_implementation blocks as ownership with exactly one
worker spawn, no escalation, targeted owned-child cleanup, working
explicit cancel, and unknown processes untouched. Keeps the injected
rc125/rc124/rc1 classification coverage. No production change.
@lukemaj
lukemaj marked this pull request as ready for review September 24, 2026 21:19
@lukemaj
lukemaj merged commit 0626db5 into main Sep 24, 2026
1 check passed
@lukemaj
lukemaj deleted the fix/88-no-elapsed-worker-kill branch September 24, 2026 21:19
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant