Skip to content

fix(orchestration): recover send after idle-released Nexplore session - #413

Merged
johnnyelwailer merged 2 commits into
mainfrom
fix/nexplore-session-timeout-recovery
Oct 7, 2026
Merged

johnnyelwailer merged 2 commits into
mainfrom
fix/nexplore-session-timeout-recovery

Conversation

@johnnyelwailer

Copy link
Copy Markdown
Owner

Summary

Phil reported two critical Nexplore failures after sync:

  1. Send after timeout → Provider session provider-session:provider-instance:nexplore:thread:… is not active.
  2. Stop that session → raw SQLite unable to open database file

Do not merge until Phil says go.

Root cause

(1) Send-to-inactive

ProviderSessionManager idle-releases live sessions (default 30m) and marks them stopped, but a projected turn can still look running. Client/server then choose steer_active / restart_active. dispatchSteerIntoRun does providerSessions.get → none → hard fail. No revive path.

Nexplore declares supportsActiveSteering: true, so auto mode prefers steer whenever a running turn is projected — making this especially visible on Nexplore after timeout.

(2) Stop + sqlite

Stop with a dead session already tries to settle locally. The raw unable to open database file is SQLITE_CANTOPEN, which on a nearly full Mac volume (~0.1 GiB free earlier today) commonly means WAL/SHM cannot be created — i.e. disk full / storage pressure, not a wrong home path. That said, the error was not actionable.

Fix

  • Steer/restart into a dead live session: settle the zombie turn (same settlement as Stop-with-no-session) + fall through to start_immediately so the user message opens a fresh session.
  • promote-to-steer safety net: settle zombie + actionable “send again” error (clears stuck state).
  • Persistence errors: detect SQLITE_CANTOPEN / SQLITE_FULL / ENOSPC / “unable to open database file” and rewrite to a clear free-disk message on EventSinkWriteError and dispatch causes.

Tests

  • persistenceStorageError.test.ts — detection + rewrite
  • DeadSessionSendRecovery.test.ts — steer into zombie running turn recovers (interrupted + fresh run); EventSink message copy
  • Focused run: 5 passed

Prior e2e QA gap (honest)

e2e-comp covered live Instant new/continue echo, partial stop-while-streaming, pack smoke/recipes, cutover. It did not cover: idle-timeout → inactive, send-after-timeout, stop-after-timeout, session revive. Checklist added under /workspace/t3sync/PROGRESS.md and e2e-comp/REPORT.md.

Notes

Phil hit two Nexplore dead-ends after sync: send into a timed-out session
failed with "Provider session … is not active", and Stop then surfaced raw
SQLite "unable to open database file" (often disk-full / WAL on a nearly
full Mac volume).

- When steer/restart targets a projected running turn whose live session is
  gone, settle the zombie turn and continue as start_immediately instead of
  dead-ending.
- Steer safety net (promote-to-steer): settle the zombie and return an
  actionable "send again" error.
- Rewrite SQLITE_CANTOPEN / ENOSPC-style persistence failures to a clear
  free-disk message on EventSink writes and dispatch error causes.
- Regression tests for storage-error detection and dead-session send recovery.

Do not merge without Phil go.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T06:42:00.477737Z 93932a4 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Oct 7, 2026
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

⚠️ The latest CI run did not produce a thread transfer result for 37e61c6.

This comment will update automatically after the next completed run.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 93932a40fc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

throughRunOrdinal: targetRun.ordinal,
now,
});
return yield* new OrchestratorDispatchError({

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Commit zombie settlement before rejecting steer promotion

When queued-message.promote-to-steer encounters a missing session, the calls above append settlement events and then this return fails dispatchOnce; dispatchWithReceiptEffect consequently commits only a rejected receipt and discards every accumulated event. The zombie run and queued message therefore remain unchanged despite the error claiming the stuck turn was cleared, so repeating the promotion still hits the same dead session. Commit the settlement as a successful state transition or perform it through a separately committed command before rejecting the promotion.

AGENTS.md reference: AGENTS.md:L43-L47

Useful? React with 👍 / 👎.

*/

const STORAGE_EXHAUSTED_PATTERN =
/unable to open database file|SQLITE_CANTOPEN|SQLITE_FULL|no space left|ENOSPC|disk(?:\s|_)?full|database (?:or disk is full|disk image is malformed)/i;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Do not classify database corruption as disk exhaustion

SQLite's database disk image is malformed message indicates database corruption, not exhausted storage. Matching it here causes rewritePersistenceFailureCause to replace the diagnostic throughout the event sink with advice to free disk space and retry or restart, concealing the corruption signal and directing users toward a recovery that cannot repair it. Remove this phrase from the storage-exhaustion pattern or give corruption its own actionable error.

Useful? React with 👍 / 👎.

…hitelist

- Include runtimeRequests on dispatchSteerIntoRun projection (settleInterruptedRun).
- Un-export internal persistenceFailureText (knip unused export).
- Whitelist EventSink.ts + new persistence/DeadSession files in additive guard.
@johnnyelwailer
johnnyelwailer merged commit 556812d into main Oct 7, 2026
20 of 21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant