Skip to content
This repository was archived by the owner on Oct 1, 2026. It is now read-only.
This repository was archived by the owner on Oct 1, 2026. It is now read-only.

Recover jobs through the installed runtime after a plugin update #75

Description

@lukemaj

Outcome

A recoverable failure caused by a removed plugin runtime can continue through the currently installed compatible Model Router runtime, preserving the job and its work.

This advances the Objective of durable dispatcher-owned delivery. Updated direction from the user on 2026-09-24: prefer recovery onto the installed runtime over keeping permanent per-job runtime copies for a rare event.

Evidence

The original ao-core-4 observation did not prove a spawn failure: the correction in this issue's comments records two later successful turns after the old directory disappeared. Preserve that distinction.

A later occurrence during toolboxmd/chromeria#1 left t3-fleet-pane-1-v2 blocked with codex_resume_failed rc=127: turn not completed after its old plugin-cache path was removed. Temporary restoration of the old path was a workaround, not the desired product behavior. See the T3 issue's durable handoff for the recovery record.

Acceptance criteria

  • A missing old runtime path is diagnosed with a concrete cause and next action rather than leaving an opaque sticky failure.
  • Recovery resolves the currently installed runtime and checks whether it can read the stored job/invocation state. Compatible recovery preserves the logical job, task, candidate, session identities and pending/completed actions.
  • Establish old controller/worker ownership and termination before transferring execution. Never create a duplicate writer or release an uncertain workspace claim. Work that is still healthy is not interrupted merely because an update occurred.
  • Record the actual runtime and policy used before and after recovery. An unsupported state or policy transition reports the specific incompatibility and retains work; it must not silently change an explicit requested route or repeat an already-applied action.
  • The dispatcher can continue through the supported recovery path without asking the planner to restore cache directories or implement a fix inside the active product task.

Simplification decision

Replace the previous requirement to copy every runtime permanently and prune old copies. Do not add a runtime archive, updater service or new daemon. Reuse installed-runtime discovery, durable job state and existing recovery/ownership checks.

Non-goals

Changing marketplace installation, silently installing a different version, unbounded retries, irreversible state migrations, or a blanket guarantee that arbitrary future runtime versions are compatible.

Blockers

None. Related malformed-envelope recovery remains in #77; broader dispatcher recovery and telemetry have a separate owning issue.

Required proof

Use disposable package copies and isolated state, never the real installed cache. Reproduce removal of an old runtime between turns; prove compatible installed-runtime recovery, no duplicate writer/action, preserved identity and truthful version evidence. Cover a healthy surviving child, unknown ownership and an incompatible replacement. Run the project's applicable complete proof and independent review. Record ordinary-work evidence when available; do not claim universal upgrade compatibility from the fixture.

Related recovery spec: #87

Architecture audit refinement (2026-09-24)

The runtime-recovery proof must distinguish an action that provably never started from an action whose effects are unknown. Today a supervisor spawn error can be persisted as failed rc127 and permanently memoized by the action key. Merely restoring or locating a runtime does not make that action retryable. Make never-started failures recoverable with explicit evidence; never replay live, successful, or uncertain actions.

Stored controller/meta JSON also needs a concrete compatibility decision at the recovery boundary; additive database migrations alone are not a blanket cross-version compatibility guarantee. Use the smallest explicit supported compatibility check, retain an unsupported-state reason, and preserve policy/route identity. Do not create a general migration framework or runtime archive.

The T3 rc127 observation remains evidence of failure after runtime removal; empty output alone cannot prove the exact spawn-error branch. Prove that branch with disposable package copies and isolated state. Preserve productive workers during upgrades and establish process ownership before transferring control.

Activity

  1. lukemaj commented on Sep 23, 2026

    @lukemaj
    ContributorAuthor

    Correction to the problem statement: job ao-core-4 completed two further turns after the 0.26.1 directory was deleted (an opencode_control of 183 s and a codex_resume of 134 s, both spawned after 22:17Z) and succeeded. So a spawn after the deletion did not fail in this case, and the claim that the next spawn cannot find the runner code is unconfirmed. The risk stands: a running job's code directory can vanish under it, and its behavior then depends on how each child resolves the runner package. The deterministic deletion test in the acceptance criteria should establish the actual behavior before choosing the fix.

    🤖 Generated with Claude Code

  2. changed the title [-]Jobs survive a plugin upgrade: run from a runner copy the updater never deletes[/-] [+]Recover jobs through the installed runtime after a plugin update[/+] on Sep 24, 2026
  3. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Implementation is running in job router-75-runtime-recovery, with its own Luna dispatcher and exclusive branch fix/75-compatible-runtime-recovery from 7e64edfe56aae9290b01bdbb4a69ad0dddaecd51. Observer task model-router-75. Requested hard lane; selected Muse xhigh on Go after recorded free-pool exhaustion. Observed worker execution will be reported separately from route selection.

    This job uses a fixed task-local development runtime from reviewed #88 candidate 77d0010371fa36e2ee77567032c8fa3f3cafa6fc (0.29.2, policy2.7.1), because the installed0.29.1 still contains the elapsed-time killer. Independent review confirmed no age-based worker termination or duplicate-writer path in this candidate, while requesting two recovery edge-case corrections now assigned to router-88-review-fixes. It is not an installed or released artifact. All three new dispatch invocations store timeout_secs: null.

    Authority: scoped implementation, full proof, commits, branch push and one draft PR to main. No merge/release/install. Preserve all other workspaces and live native records. Exact candidate, complete proof and independent final review are pending. Router75 must reconcile its final integration with #88 before delivery.

  4. lukemaj commented on Sep 24, 2026

    @lukemaj
    ContributorAuthor

    Merged PR97 at823f261241427862d142efacf86f413b63d96078 after full615test exact-head CI and independent Opus approval on d74fbe1. Includes merged88 with no elapsed kill regression. Remaining small recovery handoff caveats are passed to87; installation and universal upgrade compatibility are not claimed. Task workspace retained while dependent87 implementation and runtime provenance references need it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions