Pilot implementation: #282
Evaluate an upstream proposal only after the LastCode recovery pilot proves useful on a real recurrence.
The pilot detects failed Codex/Claude event recording, reconciles provider-confirmed finished turns using the existing liveness cadence, and offers a separate repair conversation when deterministic recovery cannot finish. It preserves queued messages and does not infer a dead run from inactivity.
Before preparing an upstream PR:
- Record a real incident where the feature detected or repaired a stuck thread, including the failure mode and outcome.
- Confirm that healthy long-running turns are unaffected and no original work was repeated.
- Identify the smallest generally useful change against current upstream behavior, accounting for upstream fixes already merged or in flight.
- Remove fork-specific repair-agent assumptions unless upstream explicitly wants them.
If the problem does not recur, leave this deferred and do not prepare an upstream PR. Automated tests alone do not satisfy the real-world usefulness gate.
Pilot implementation: #282
Evaluate an upstream proposal only after the LastCode recovery pilot proves useful on a real recurrence.
The pilot detects failed Codex/Claude event recording, reconciles provider-confirmed finished turns using the existing liveness cadence, and offers a separate repair conversation when deterministic recovery cannot finish. It preserves queued messages and does not infer a dead run from inactivity.
Before preparing an upstream PR:
If the problem does not recur, leave this deferred and do not prepare an upstream PR. Automated tests alone do not satisfy the real-world usefulness gate.