Skip to content
This repository was archived by the owner on Oct 1, 2026. It is now read-only.
This repository was archived by the owner on Oct 1, 2026. It is now read-only.

Report terminal job state back to the planner: ready to merge or blocked #94

Description

@lukemaj

Outcome

When a job reaches a terminal state, the runner wakes the saved planner in its own harness with the result, and the planner reports back in its thread: ready to merge (with the PR URL) or blocked / failed / cancelled (with the reason). Nobody has to poll status to learn a job ended.

Problem

_complete_job (runner/controller.py) persists the result and PR URL to the ledger and result file, and nothing else. Blocked, failed and cancelled jobs are equally silent. The planner is woken only for planner_question. The worker opens the PR under the user's own gh login, so GitHub does not notify the user either. Blocked jobs (for example the three #77 occurrences overnight) wait unnoticed until someone runs status.

Acceptance criteria

  • On every terminal state (succeeded, blocked, failed, cancelled) the runner sends one terminal report to the saved planner session through the existing callback path (Claude resume, Codex exec resume, OpenCode run --session, Grok fresh session with the handoff summary).
  • The report carries request id, status, PR URL when present, block or failure reason, and the handoff summary.
  • The terminal report is sent after the terminal state is persisted and never changes the job's result or status.
  • A busy planner (a running process holding the session) does not lose the report: delivery is retried until the planner is idle or a bounded delivery window ends; the outcome of delivery is recorded on the job and visible in status.
  • Exactly one delivered report per terminal state; recover and restarted controllers do not send duplicates.
  • A test per harness adapter with fakes, plus a regression for the busy-planner case.

Non-goals

  • Push notifications (Bark or other) and a configurable notify command. Possible follow-up after measured need.
  • Notifications for intermediate states.

Blockers

Overlaps the blocking and timeout paths edited by #88 (PR #90). Start after #90 merges.

Proof

python3 scripts/test.py, plus one live job whose completion is reported back in a Claude Code planner thread.

Objective contribution

Milestone item, user-confirmed 2026-09-24. The Objective ends each job at one open PR for a human merge and keeps the human off the critical path; a silent PR or a silently blocked job defeats both.

Activity

  1. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Findings from reading the callback path, 2026-09-24:

    • The dispatcher's only way to reach a Claude planner is to start a second headless process on that session (claude --resume SID -p). Nothing can inject a message into a thread the user has open.
    • Busy detection misses T3 threads. adapters.planner_session_in_use matches the session id only as a separate argv token. T3's live Claude process passes --session-id=<SID> as one token, so the check never matches and the open thread is not seen as busy. Only wait_planner_quiet (transcript still changing) blocks.
    • Inference, not yet verified: when the thread is open but idle, the headless resume runs and appends its exchange to the transcript file. The live T3 process keeps its in-memory history, so the report likely never appears in the open window, and the thread's next turn may continue from its own last message and skip the report.

    So for #94 a retried report is not enough when the thread is open. Add to scope: verify the open-idle behavior with a live probe; fix the --session-id= match; and choose a delivery path that reaches an open thread, which is a push into the live session (as #86 does for Codex through app-server) rather than a second process.

  2. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Scope: live delivery must cover every supported planner harness, not only Claude. User-confirmed 2026-09-24. What is known per harness, from local CLI help on this Mac (OpenCode 1.18.32, Grok Build 1.0.41); none of it is probed yet:

    Planner harness Candidate path into an open session Status
    Codex app-server steering of the same thread Proven by #86's live probe
    OpenCode the TUI is a client of an opencode serve server (opencode attach <url>); posting a message to the session through that server should reach the open TUI Inference; the runner must also find the user's server, not start its own
    Grok Build grok --resume <SESSION_ID> exists, so the current "no usable resume, answer from a fresh session" fallback looks outdated; running TUIs attach to a leader process on ~/.grok/leader.sock Unknown whether the leader accepts injected messages
    Claude Code in T3 T3 owns the live process's stdin Unknown; needs a host-side injection path

    Add to acceptance: one live probe per harness showing a report reaching an open thread, or a recorded reason that harness cannot, with its fallback. Re-check Grok resume and switch callbacks from the fresh-session fallback to real resume if it holds.

  3. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Lead for the Claude Code row: a live Claude Code session exports CLAUDE_CODE_MESSAGING_SOCKET (a per-process Unix socket, e.g. /tmp/cc-socks/<pid>.sock) and CLAUDE_CODE_MESSAGING_TOKEN, and Claude Code offers messaging to other local Claude sessions (the SendMessage and ListAgents tools list them). That is a candidate path for injecting a report into an open Claude thread, including in T3. Unverified: the socket protocol and whether a message reaches an idle session's next turn. Probe it before choosing a delivery path.

  4. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Shared capability with toolboxmd/chromeria#1 (Fleet pane): both need to deliver a message into an existing live native session for Codex, Claude Code, OpenCode and Grok. #86 already provides the Codex half for the dispatcher. Reuse or contribute the per-harness live-delivery probes with t3code#1 instead of building a second, Router-only injection path; the Router side should only decide when to report and what, and use the same native mechanism T3 uses.

  5. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Scope narrowed for now (2026-09-24): build the runner-side trigger (one report per terminal state, with status, PR URL or reason) through the existing callback path, plus the --session-id= busy-match fix. Hold the live-delivery transport until toolboxmd/chromeria#3 decides whether T3 spawns and owns the sessions, in which case delivery becomes a T3 thread message.

  6. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Submitted as Router job router-94-terminal-report (installed 0.29.1, lane hard, worktree model-router-worktrees/terminal-94, branch feat/94-terminal-report) with the narrowed scope: runner-side end-of-job trigger through the existing callback path plus the --session-id= busy-match fix. No live-delivery transport; toolboxmd/chromeria#3 found T3 can own that. PR will follow; not merged.

  7. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Runner-side trigger done: PR #99 (branch feat/94-terminal-report, 0.30.0).

    On every terminal state the runner persists first, then sends exactly one report through the existing callback path with request id, status, PR URL or reason, and handoff summary (succeeded says ready to merge). Busy planners retry within a bounded window; the outcome is recorded on the job and visible in status; recover sends no duplicates. Includes the --session-id=/--resume= busy-match fix. No live-delivery transport, per the narrowed scope.

    Proof: python3 scripts/test.py — 613 tests OK (1 skipped), plus a live fake-CLI loop showing question callback + terminal report with no duplicate on recover. Not merged; needs human review/merge.

  8. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Job router-94-terminal-report blocked at 01:00 with implementation_failed rc=124: the correction turn hit the old 1800-second worker cap that #88 removed on main but the installed 0.29.1 still has. Partial correction kept as a local checkpoint commit. Continued as router-94-terminal-report-2 on the dispatcher's correction instructions, run from a source checkout of main at 823f261 (not an install) with its own state dir ~/.local/share/durable-runner-planner-f1213, so it has the #88 fix and does not touch the shared ledger. Same branch and PR #99.

  9. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Correction pushed for #94 on PR #99 at e8257ed (rebased on origin/main 823f261, VERSION 0.31.0).

    What changed vs 1a31f58:

    • Bounded busy retry for every resume harness via planner_process_in_use (Claude/Codex/OpenCode, Grok fresh-session exempt); keeps --session-id=/--resume=/--session= matching.
    • Durable per-event terminal_report send under a stable action key keyed by the terminal ledger event id; finished sends adopted without resending, live sends deferred, delivered claimed only after durable record, record failures visible without changing job status/result; planner_question rows never adopted.
    • CLI cancel/recover use the durable path; step and controller tail still report after persist.

    Proof:

    • python3 -m unittest tests.test_terminal_report94: 26 tests OK.
    • python3 -m compileall -q runner tests scripts: compile-ok.
    • python3 scripts/test.py: 640/641 OK in 345s; single failure test_cli_conflict_clears_via_questions_clear (live-controller timing) passes in isolation (1 test OK).
    • versionctl doctor: OK; release-check: v0.31.0 at e8257ed.

    PR: #99 (single open PR, base main, head e8257ed). No merge, no live transport, no push notifications.

  10. lukemaj commented on Sep 25, 2026

    @lukemaj
    ContributorAuthor

    Job router-94-terminal-report-2 succeeded: PR #99 at e8257ed (0.31.0). Dispatcher correction complete: durable per-event report send with crash recovery and no duplicates, bounded busy retry on every planner harness, delivery outcome visible in status, all terminal paths covered, --session-id=/--resume= busy-match fix kept. Worker proof 641 tests OK (1 skipped), 26 targeted terminal-report tests, CI proof pass. Scope stays runner-side only; no live-delivery transport. Ready for human merge; #98 (0.30.0) and #99 both touch the callback path, so merge #98 first and rebase #99.

  11. lukemaj commented on Sep 25, 2026

    @lukemaj
    ContributorAuthor

    Done: PR #99 (0.33.0) sends one end-of-job report per terminal state to the saved planner, and the T3 path (#106, PR #109, 0.34.0) posts it into the planner's T3 thread. Closing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions