Outcome
A recoverable failure caused by a removed plugin runtime can continue through the currently installed compatible Model Router runtime, preserving the job and its work.
This advances the Objective of durable dispatcher-owned delivery. Updated direction from the user on 2026-09-24: prefer recovery onto the installed runtime over keeping permanent per-job runtime copies for a rare event.
Evidence
The original ao-core-4 observation did not prove a spawn failure: the correction in this issue's comments records two later successful turns after the old directory disappeared. Preserve that distinction.
A later occurrence during toolboxmd/chromeria#1 left t3-fleet-pane-1-v2 blocked with codex_resume_failed rc=127: turn not completed after its old plugin-cache path was removed. Temporary restoration of the old path was a workaround, not the desired product behavior. See the T3 issue's durable handoff for the recovery record.
Acceptance criteria
Simplification decision
Replace the previous requirement to copy every runtime permanently and prune old copies. Do not add a runtime archive, updater service or new daemon. Reuse installed-runtime discovery, durable job state and existing recovery/ownership checks.
Non-goals
Changing marketplace installation, silently installing a different version, unbounded retries, irreversible state migrations, or a blanket guarantee that arbitrary future runtime versions are compatible.
Blockers
None. Related malformed-envelope recovery remains in #77; broader dispatcher recovery and telemetry have a separate owning issue.
Required proof
Use disposable package copies and isolated state, never the real installed cache. Reproduce removal of an old runtime between turns; prove compatible installed-runtime recovery, no duplicate writer/action, preserved identity and truthful version evidence. Cover a healthy surviving child, unknown ownership and an incompatible replacement. Run the project's applicable complete proof and independent review. Record ordinary-work evidence when available; do not claim universal upgrade compatibility from the fixture.
Related recovery spec: #87
Architecture audit refinement (2026-09-24)
The runtime-recovery proof must distinguish an action that provably never started from an action whose effects are unknown. Today a supervisor spawn error can be persisted as failed rc127 and permanently memoized by the action key. Merely restoring or locating a runtime does not make that action retryable. Make never-started failures recoverable with explicit evidence; never replay live, successful, or uncertain actions.
Stored controller/meta JSON also needs a concrete compatibility decision at the recovery boundary; additive database migrations alone are not a blanket cross-version compatibility guarantee. Use the smallest explicit supported compatibility check, retain an unsupported-state reason, and preserve policy/route identity. Do not create a general migration framework or runtime archive.
The T3 rc127 observation remains evidence of failure after runtime removal; empty output alone cannot prove the exact spawn-error branch. Prove that branch with disposable package copies and isolated state. Preserve productive workers during upgrades and establish process ownership before transferring control.
Outcome
A recoverable failure caused by a removed plugin runtime can continue through the currently installed compatible Model Router runtime, preserving the job and its work.
This advances the Objective of durable dispatcher-owned delivery. Updated direction from the user on 2026-09-24: prefer recovery onto the installed runtime over keeping permanent per-job runtime copies for a rare event.
Evidence
The original ao-core-4 observation did not prove a spawn failure: the correction in this issue's comments records two later successful turns after the old directory disappeared. Preserve that distinction.
A later occurrence during toolboxmd/chromeria#1 left t3-fleet-pane-1-v2 blocked with
codex_resume_failed rc=127: turn not completedafter its old plugin-cache path was removed. Temporary restoration of the old path was a workaround, not the desired product behavior. See the T3 issue's durable handoff for the recovery record.Acceptance criteria
Simplification decision
Replace the previous requirement to copy every runtime permanently and prune old copies. Do not add a runtime archive, updater service or new daemon. Reuse installed-runtime discovery, durable job state and existing recovery/ownership checks.
Non-goals
Changing marketplace installation, silently installing a different version, unbounded retries, irreversible state migrations, or a blanket guarantee that arbitrary future runtime versions are compatible.
Blockers
None. Related malformed-envelope recovery remains in #77; broader dispatcher recovery and telemetry have a separate owning issue.
Required proof
Use disposable package copies and isolated state, never the real installed cache. Reproduce removal of an old runtime between turns; prove compatible installed-runtime recovery, no duplicate writer/action, preserved identity and truthful version evidence. Cover a healthy surviving child, unknown ownership and an incompatible replacement. Run the project's applicable complete proof and independent review. Record ordinary-work evidence when available; do not claim universal upgrade compatibility from the fixture.
Related recovery spec: #87
Architecture audit refinement (2026-09-24)
The runtime-recovery proof must distinguish an action that provably never started from an action whose effects are unknown. Today a supervisor spawn error can be persisted as failed rc127 and permanently memoized by the action key. Merely restoring or locating a runtime does not make that action retryable. Make never-started failures recoverable with explicit evidence; never replay live, successful, or uncertain actions.
Stored controller/meta JSON also needs a concrete compatibility decision at the recovery boundary; additive database migrations alone are not a blanket cross-version compatibility guarantee. Use the smallest explicit supported compatibility check, retain an unsupported-state reason, and preserve policy/route identity. Do not create a general migration framework or runtime archive.
The T3 rc127 observation remains evidence of failure after runtime removal; empty output alone cannot prove the exact spawn-error branch. Prove that branch with disposable package copies and isolated state. Preserve productive workers during upgrades and establish process ownership before transferring control.