Repository navigation
Remove elapsed-time termination of active agents and jobs #88
Description
Activity
Issue88 initial implementation is PR #90 at
77d0010371fa36e2ee77567032c8fa3f3cafa6fc, version0.29.2, policy2.7.1. Exact-candidate GitHub deterministic proof passed; local final suite594 tests passed with1skip. Initial jobrouter-88-no-elapsed-killsucceeded and all its invocations ended. This is execution success, not acceptance.Independent Opus5.5 high review requested two corrections: idle OpenCode sessions with incomplete terminal output can evade existing stall handling; supervisor-loss rc125 must retain ownership handling on the Grok path. Age-kill removal and duplicate-writer prevention were checked and cleared. No new absolute deadline is requested.
Source ownership now belongs exclusively to correction job
router-88-review-fixeson the same branch/workspace and existing PR90. Dispatcher owns correction; planner remains read-only for implementation. Observer taskmodel-router-88. Full proof and independent re-review required on the changed candidate. No second version bump for these pre-delivery corrections.The three jobs88 correction,75 runtime recovery and Observer25 run through a fixed task-local snapshot of reviewed77d0010 so productive agents are not terminated at30minutes by installed0.29.1. All three live dispatch records confirm policy2.7.1 and null timeout. This is a development runtime, not an installed/released artifact.
Authority remains implementation, proof, push and existing draftPR updates; no merge/release/install. Final acceptance and delivery are unfinished. Conditional inspector #89 remains deferred.
Canonical handoff: #88 implementation complete, review-ready
Outcome: productive agents remain running regardless of total elapsed time. Elapsed-time killing is deleted, not increased: no per-turn or per-job deadline exists in the runner.
Branch/workspace/base/HEAD
- Branch:
fix/88-no-elapsed-worker-kill(exclusive worktree/Users/lukaszmaj/dev/toolboxmd/model-router-worktrees/no-kill-88, no other checkout touched, no installed-plugin edits, no persistent host service) - Base:
7e64edfe56aae9290b01bdbb4a69ad0dddaecd51(main-line); HEAD:77d0010371fa36e2ee77567032c8fa3f3cafa6fc - Pushed; draft PR: Delete elapsed-time termination of agent turns and jobs (#88) #90 (targets main, no merge)
Changed semantics / compatibility
- Agent turns (codex dispatch/resume, claude/codex/opencode/grok callbacks, opencode worker control, grok worker control) carry no elapsed deadline; recovery never cancels/drains by job age. Legacy positive
timeout_secson old rows stays readable but is never enforced. --timeout-secs/timeoutargs accepted and recorded (new invocation rows store the passed value, default NULL), but bound nothing. Documented in CLI help, RUNNER.md, policy 2.7.1 (skill reference regenerated, not hand-edited), GLOSSARY Silence window (outer-budget wording removed).- Preserved: explicit cancellation (cancel -> drain -> cancelled), real terminal failures, process-ownership/unknown-owner safeguards, stream-stall handling (policy windows pass through unclamped), bounded transport/readiness/cleanup limits (probe 30s, serve startup/health 120s, prompt-accept 60s, wait_idle 30/10s, SIGTERM/SIGKILL drain bounds, supervisor reap 10s, proof 600s). Legacy
cancel_requested=2drains still finalize with timeout origin. rc124/143 controller paths untouched.
Proof (exact candidate 77d0010)
- Before (runner/ stashed): 7/9 new tests fail on old code (rc124 kill, 1800 defaults, clamped window, recover timeout action) — reproduces the old kill deterministically in seconds.
- After:
python3 scripts/test.py-> 594 tests OK (1 skipped) in ~316s;python3 -m compileall -q runner scripts testsclean;python3 -m runner.policy validateclean. Full logs retained in worker notes. - Updated tests that encoded the rejected deadline (with Remove elapsed-time termination of active agents and jobs #88 justification, stall/cancel/ownership/failure regressions kept strict): test_issue73 bound + boundary classes, test_faults hung-turn + job-timeout (now: legacy timeout leaves active job running, cancel stops it), test_grok hung worker (pinned stall_secs), test_stall33 header.
Version: 0.29.1 -> 0.29.2 (one patch via versionctl, no release). release-check notes VERSION_CONFLICT vs tag v0.29.0 only because 0.29.1 was committed untagged; branch history shows the single 0.29.1 -> 0.29.2 step. No tag created (release is out of scope).
Review/delivery state: independent exact-candidate review NOT yet done — required before merge. Draft PR only; no merge, release, or install performed (no authority for those).
Remaining limitations / next action: dispatcher to arrange independent exact-candidate review of 77d0010 and (separately authorized) merge. Agent Observer measurement gap: no installed Observer skill/capture was available in this worker, so evidence availability was verified only by the ledger/measurement assertions in-suite (longest_silence_secs, elapsed_secs, terminal_class); no monitoring was added per non-goal. Future inspector work stays conditional issue #89.
- Branch:
- added a commit that references this issue
on Sep 24, 2026 Handoff: review corrections delivered (dispatcher)
Candidate: da15150, branch fix/88-no-elapsed-worker-kill, PR #90 (OPEN, head now da15150). Same outcome, same PR, no merge/install/release.
Adjudication: both independent-review findings on 77d0010 accepted and corrected; complete removal of elapsed-time killing preserved. No inspector, scheduler, watchdog, monitoring service, recovery redesign, or absolute deadline added.
Finding 1 (OpenCode idle-incomplete, medium): confirmed reachable. runner/supervisor.py non-busy branch now applies the existing tracker + stall/probe/abort/idle-confirm machinery when the turn is not terminal and not prompt-waiting (condition seen_active or new). Live part/reasoning/tool activity keeps the turn alive; only genuine silence past the unclamped policy window stalls, with the existing stalled signal/evidence shape plus idle_without_terminal_result marker. Prompt-accept 60s readiness bound untouched.
Finding 2 (Grok rc125 supervisor loss, low-medium): confirmed reachable. runner/controller.py Grok check is now rc in (124, 143, 125) with timeout language removed; rc125 blocks as ownership (report left, job blocked, no escalation behind the unsupervised child). OpenCode no-control-result path and all other rc behavior unchanged.Changed files: runner/supervisor.py, runner/controller.py, tests/fakes.py (idle_incomplete/idle_empty modes), tests/test_issue88.py (+4 tests). No VERSION/CHANGELOG/mirror/policy change: 0.29.2 from 0.29.1 and policy 2.7.1 preserved. Note: the pre-commit version gate was bypassed (--no-verify) for this commit only, per the explicit no-repeat-bump instruction; message conventions kept, no other hook behavior skipped.
Proof on da15150 (exact candidate; 77d0010 proof not reused):
- python3 -m unittest tests.test_issue88 -> 13 tests OK (~25s)
- python3 -m unittest tests.test_grok tests.test_stall33 -> 66 tests OK (~64s)
- python3 scripts/test.py -> Ran 598 tests in 329.247s, OK (skipped=1)
- python3 -m compileall -q runner scripts tests -> clean
- Coverage: idle-incomplete and idle-empty stall with existing stall evidence/controller path (stalled_retry, stream_silence, idle_confirmed, report + terminal_class stalled); rc125 blocked-as-ownership with single-spawn, unknown-process-alive, cancel-works, and rc1-still-fails assertions; prior Remove elapsed-time termination of active agents and jobs #88 regressions (active work past legacy caps, legacy timeout readable/unenforced, NULL defaults, recovery adopt, silence, cancel, ownership) all still green.
Delivery state: correction committed and pushed; PR #90 updated; independent exact-candidate re-review remains coordinator-owned. No unchecked coverage beyond re-review itself. No blockers.
- added a commit that references this issue
on Sep 24, 2026 Handoff: faithful supervisor-loss regression delivered (dispatcher)
Candidate: 3a24647, branch fix/88-no-elapsed-worker-kill, PR #90 (OPEN). Same outcome, same PR, no merge/install/release. Production fixes from da15150 unchanged and accepted as-is.
What changed vs da15150 (tests only): tests/test_issue88.py gains test_supervisor_killed_with_live_child_returns_125_and_blocks, the faithful reviewed trigger: a real core._durable_run (thread, fake Grok hold child, make_durable_run_cmd) loses its supervisor via SIGKILL on the exact recorded supervisor pid only. The durable run returns rc 125 with the owned child live and _invocation_ownership live; that authentic result through controller.run_implementation blocks as ownership (implementation_failed reason), with exactly one worker spawn ever, one controller call, no escalation, targeted owned-child cleanup by exact pid, working explicit cancel, and an unrelated sleep process surviving throughout. A spawn gate (worker logged before the kill) fixed a log-read race the full suite exposed under load. Kept the injected rc125/rc124/rc1 classification tests as distinct fast coverage. Mutation-checked: reverting the controller fix makes the drill fail with implementation_failed instead of blocked.
Proof on 3a24647 (exact candidate; earlier proof not reused):
- python3 -m unittest tests.test_issue88 -> 14 tests OK (~26s)
- python3 -m unittest tests.test_grok tests.test_stall33 -> 66 tests OK
- Faithful drill: stable across repeats and alongside GrokDurable/GrokControllerInjected; no stray processes (verified unrelated durable-runner supervisors untouched)
- python3 scripts/test.py -> Ran 599 tests in 334.931s, OK (skipped=1), log-verified
- python3 -m compileall -q runner scripts tests -> clean
- Version 0.29.2 / policy 2.7.1 preserved, no VERSION/CHANGELOG/mirror change; pre-commit version gate bypassed (--no-verify) for correction commits only per the explicit no-repeat-bump instruction.
Delivery state: committed, pushed, PR #90 updated. Independent exact-candidate re-review remains coordinator-owned. No blockers.
- added a commit that references this issue
on Sep 24, 2026 Resolved by merged PR90, merge commit0626db54928636bb1f0b4aa5fdff328ee2f928aa. Independent Opus5.5 high re-review approved exact candidate3a24647aa57f0cff889302a0a55d70a3771f5a2b after both findings were fixed. Exact-candidate full CI passed (run36055751314): no elapsed-time worker/job killing, legacy budgets ignored, explicit cancellation/stall/ownership preserved, faithful supervisor-loss regression included. Merge authorized by user in the ongoing overnight task. Installation is not claimed; dependent runtime-recovery work remains in75/PR97. Retain task worktree and fixed task runtime while dependent jobs need them.
Parent
#87
Outcome
Productive agents remain running regardless of total elapsed time. Model Router does not stop an agent or drain a job merely because an invocation or job age crosses a fixed budget.
Contribution: preserve dispatcher ownership through completion, advancing the current Objective of delegated work reaching a verified open PR without human recovery.
Evidence and Elon decision
Job
t3-fleet-router-86-complete, runner 0.29.1, invocation2175332f48e87ebc: requested job budget 5400 seconds, persisted worker cap 1800 seconds, terminated after 1801.131 seconds with rc124 andimplementation turn timed out. Longest recorded silence was 12.77 seconds; proof exit was 0. Activity does not establish useful productivity, but no stall justified this elapsed-time stop. See #87 reproduction comment.The inherited premise that every agent needs an absolute time budget is rejected by the user. Delete elapsed-time killing instead of increasing the constant. Keep explicit cancellation, confirmed failure handling, process ownership, and existing stream-stall handling. Agent Observer owns measurements and evidence of unusually long or stalled work. Do not add an inspector, scheduler, timeout replacement, or new monitoring service. The present constraint is the runner stopping active work; one bounded patch and clock-controlled regression proof is the smallest surviving solution.
Acceptance criteria
Non-goals
Blocked by
None. Authorized for implementation now. Independent isolated branch from main; preserve the dirty #86 workspace.
Required proof
Use deterministic clocks and faithful subprocess/protocol fixtures to reproduce the old kill and prove continued active work beyond both old invocation and job deadlines. Include job recovery with legacy timeout records, normal completion and cleanup, explicit cancellation, genuine stream silence, and unknown ownership. Do not spend 30 minutes reproducing a constant. Run the full project suite and compile check, retain exact-candidate independent review, and open one PR. Update tests that encoded the rejected deadline requirement without weakening valid cancellation/ownership regressions.