Repository navigation
fix(orchestration): recover send after idle-released Nexplore session - #413
Conversation
Phil hit two Nexplore dead-ends after sync: send into a timed-out session failed with "Provider session … is not active", and Stop then surfaced raw SQLite "unable to open database file" (often disk-full / WAL on a nearly full Mac volume). - When steer/restart targets a projected running turn whose live session is gone, settle the zombie turn and continue as start_immediately instead of dead-ending. - Steer safety net (promote-to-steer): settle the zombie and return an actionable "send again" error. - Rewrite SQLITE_CANTOPEN / ENOSPC-style persistence failures to a clear free-disk message on EventSink writes and dispatch error causes. - Regression tests for storage-error detection and dead-session send recovery. Do not merge without Phil go.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Thread transfer impact
This comment will update automatically after the next completed run. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 93932a40fc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| throughRunOrdinal: targetRun.ordinal, | ||
| now, | ||
| }); | ||
| return yield* new OrchestratorDispatchError({ |
There was a problem hiding this comment.
Commit zombie settlement before rejecting steer promotion
When queued-message.promote-to-steer encounters a missing session, the calls above append settlement events and then this return fails dispatchOnce; dispatchWithReceiptEffect consequently commits only a rejected receipt and discards every accumulated event. The zombie run and queued message therefore remain unchanged despite the error claiming the stuck turn was cleared, so repeating the promotion still hits the same dead session. Commit the settlement as a successful state transition or perform it through a separately committed command before rejecting the promotion.
AGENTS.md reference: AGENTS.md:L43-L47
Useful? React with 👍 / 👎.
| */ | ||
|
|
||
| const STORAGE_EXHAUSTED_PATTERN = | ||
| /unable to open database file|SQLITE_CANTOPEN|SQLITE_FULL|no space left|ENOSPC|disk(?:\s|_)?full|database (?:or disk is full|disk image is malformed)/i; |
There was a problem hiding this comment.
Do not classify database corruption as disk exhaustion
SQLite's database disk image is malformed message indicates database corruption, not exhausted storage. Matching it here causes rewritePersistenceFailureCause to replace the diagnostic throughout the event sink with advice to free disk space and retry or restart, concealing the corruption signal and directing users toward a recovery that cannot repair it. Remove this phrase from the storage-exhaustion pattern or give corruption its own actionable error.
Useful? React with 👍 / 👎.
…hitelist - Include runtimeRequests on dispatchSteerIntoRun projection (settleInterruptedRun). - Un-export internal persistenceFailureText (knip unused export). - Whitelist EventSink.ts + new persistence/DeadSession files in additive guard.
Summary
Phil reported two critical Nexplore failures after sync:
Provider session provider-session:provider-instance:nexplore:thread:… is not active.unable to open database fileDo not merge until Phil says go.
Root cause
(1) Send-to-inactive
ProviderSessionManageridle-releases live sessions (default 30m) and marks themstopped, but a projected turn can still lookrunning. Client/server then choosesteer_active/restart_active.dispatchSteerIntoRundoesproviderSessions.get→ none → hard fail. No revive path.Nexplore declares
supportsActiveSteering: true, so auto mode prefers steer whenever a running turn is projected — making this especially visible on Nexplore after timeout.(2) Stop + sqlite
Stop with a dead session already tries to settle locally. The raw
unable to open database fileis SQLITE_CANTOPEN, which on a nearly full Mac volume (~0.1 GiB free earlier today) commonly means WAL/SHM cannot be created — i.e. disk full / storage pressure, not a wrong home path. That said, the error was not actionable.Fix
start_immediatelyso the user message opens a fresh session.EventSinkWriteErrorand dispatch causes.Tests
persistenceStorageError.test.ts— detection + rewriteDeadSessionSendRecovery.test.ts— steer into zombie running turn recovers (interrupted + fresh run); EventSink message copyPrior e2e QA gap (honest)
e2e-comp covered live Instant new/continue echo, partial stop-while-streaming, pack smoke/recipes, cutover. It did not cover: idle-timeout → inactive, send-after-timeout, stop-after-timeout, session revive. Checklist added under
/workspace/t3sync/PROGRESS.mdande2e-comp/REPORT.md.Notes
main(a2bdcc67, after feat(widgets): shim show_widget html onto upstream HtmlRender #408 merge) — no conflict with feat(widgets): shim show_widget html onto upstream HtmlRender #408.