Skip to content

Five bugs found while running the orchestrator v2 branch (#2829) day to day #13331

Description

@tannerpolley

I run a personal fork on top of the orchestrator v2 PR (#2829) and hit five bugs. Four are in #2829's code; one (3) is on main. I checked each against #2829's current head (3e4ca4c532) and main (effaab94e3), and they're still there. I have small, tested fixes for each and am happy to open focused PRs (against #2829's branch for 1, 2, 4, 5 and against main for 3) if you'd like them. Just say which, so you don't have to re-fix them yourselves.

1. Interrupted threads never resume after a restart (v2)

With "Continue threads after restarts" on, a thread that was mid-turn when the server restarted is cancelled and waits for a manual "continue".

restartContinuationRun() in apps/server/src/orchestration-v2/RestartContinuation.ts skips any run whose provider session status isn't "running". But only the OpenCode adapter ever writes "running"; the shared ProviderSessionManager writes ready/stopped. Everywhere else v2 treats a session as live when it's not stopped/error (Orchestrator.ts, ProjectionStore.ts). Fix: check session.status === "stopped" || session.status === "error" instead (3 lines plus a test).

2. A steered subagent shows as finished in Lineage while it runs again (v2)

After a subagent finishes and is sent new work, Lineage still shows it as completed and leaves it out of the "N running" count.

The parent's subagent record settles when the spawning run ends, and a new run on the child thread doesn't reopen it. deriveThreadRelationshipGraph() in packages/client-runtime/src/state/threadRelationships.ts uses subagent.status for the edge, and ThreadRelationshipsControl.tsx counts running agents from the records, while thread nodes already use thread.activityRunStatus. Fix: prefer the child thread's live activityRunStatus for subagent edges and the count, falling back to the record.

Related but different trigger: #7314.

3. A hung page makes preview_navigate outlive its deadline and disconnects the browser host (main)

When a page never finishes loading, preview_navigate (or preview_open reusing a tab) times out. The broker then disconnects the whole client's automation connection, and the agent's current tab is forgotten, so agents fall back to opening new tabs.

In apps/web/src/components/preview/PreviewAutomationHosts.tsx, the open (reuse) and navigate cases await bridge.navigate(...) with no limit (on desktop it settles only when webContents.loadURL does), then give waitForNavigationReadiness the full request timeout instead of what's left before hostDeadlineMs. Every other operation respects the host budget from #4685. PreviewAutomationBroker.ts treats the unanswered request as a dead connection. Fix: race the navigate call against the host deadline and pass the remaining budget to the readiness wait.

Related: #12407 (same eviction mechanism, preview_wait_for).

4. Disconnecting a Codex session leaves the thread loaded, with its MCP servers running (v2)

All Codex threads share one codex app-server. Disconnecting a thread detaches it in T3, but the native thread stays loaded in the app-server along with its MCP servers (e.g. one language server per thread or subagent). Over a day I measured about 80 idle processes and 9 GB.

ProviderSessionManager.ts's detach path for multi-thread runtimes never tells the runtime to unload the native thread; the adapter interface has no unload operation, and CodexAdapterV2.ts never sends thread/unsubscribe, although the generated client supports it. Fix: an optional unloadThread on the session runtime, implemented for Codex with thread/unsubscribe and called after detaching from a shared runtime. Verified live: 11 MCP processes went to 0.

5. A failed turn shows as "waiting" while background work is still pending (v2)

If a turn fails while subagents or background tasks are still pending, the sidebar and the notification coordinator show the thread as waiting, not failed, until the background work drains.

resolveSidebarThreadStatus in apps/web/src/components/Sidebar.logic.ts returns "waiting" for runtime.status === "idle" before checking for failure, and the runtime parks at idle while background work is pending. Fix: treat idle with a failed latest run as failed, checked before the idle branch.

Activity

  1. juliusmarminge commented on Sep 24, 2026

    @juliusmarminge
    Member

    Checked all five against t3code/codex-turn-mapping @ 3e4ca4c532 (#2829) and current main @ d4cd7d5c33 (ahead of the effaab94e3 you cited). All five are still real.

    Where they live: 1, 2, 4, and 5 are orchestrator-v2 only. Those files are not on main (and 5 is a v2 regression of behavior main already got right). 3 is on main. The preview host, desktop loadURL path, and broker are identical on this branch and main, so 3 should target main. The other four should target this v2 branch. Focused PRs as you offered are welcome; I have not reviewed the patches themselves.

    1. Restart continuation skips live non-running sessions (v2) — confirmed

    restartContinuationRun only keeps an in-flight run when session.status === "running" (apps/server/src/orchestration-v2/RestartContinuation.ts). Recovery still cancels that run either way (ProviderRuntimeRecoveryService.ts); the "running" check is what decides whether a provider-runtime.continue effect is queued. With "Continue threads after restarts" on, a Codex/Claude/Cursor/ACP thread mid-turn is cancelled and left for a manual continue.

    One correction: OpenCode is not the only writer of "running". PiAdapterV2 also emits provider_session.updated with "running" on turn start. Codex, Claude, Cursor, and ACP create the session as "ready" and never promote it. ProviderSessionManager's busy/idle tracking is in-memory only; the statuses it persists are "stopped" and "error" on release. Those adapters do mark the provider thread "active" and the provider turn "running", so the session-status check is the gate that fails.

    Treating a session as live unless it is "stopped" or "error" matches Orchestrator, ProjectionStore, ProviderSwitchService, and recovery itself. Note that OpenCode also sets the session to "waiting" while a permission or question is pending; that status would start qualifying too. Prepared continuations already bypass this check, so this is the first crash of a live turn, not the second crash of an already-admitted continuation.

    2. Steered subagent stays completed in Lineage (v2) — confirmed

    deriveThreadRelationshipGraph first builds the child edge from thread.activityRunStatus, then overwrites it. Subagent edges from projection.subagents use the same source/target/kind key and replace that status with subagent.status (packages/client-runtime/src/state/threadRelationships.ts). On the parent, ThreadRelationshipsControl then groups "Previous agents" from that edge and, whenever a projection is loaded, counts "N running" only from records with status === "running".

    A new run on the child updates the child shell's activityRunStatus. Nothing writes that back onto the parent's settled subagent row. Provider-native resume during the spawning turn can reopen the row (Codex adapter tests); a follow-up after the record has settled does not. Prefer the child shell's live activityRunStatus for the edge and the count, and fall back to the record when the shell is absent. The hover card and the screen-reader status still read agent.status off the record, so those need the same preference or the tooltip stays "completed".

    #7314 is a different bug (stuck red after a usage-limit recovery, and status frozen for the whole run). Not a duplicate.

    3. Hung preview_navigate / reused preview_open outlives the broker deadline and drops the host (main) — confirmed

    This is not a v2 bug. Same code on main @ d4cd7d5c33.

    In PreviewAutomationHosts.tsx, the open path that reuses a tab and the navigate path both await previewBridge.navigate(...) with no deadline, then pass the full request.timeoutMs (or input.timeoutMs) into waitForNavigationReadiness instead of the time left before hostDeadlineMs. On desktop, PreviewManager.navigate awaits webContents.loadURL until that load settles (apps/desktop/src/preview/Manager.ts), so a page that never finishes loading never answers. The broker treats an unanswered request as a dead connection, disconnects that client, and removeConnectionFromState deletes its assignments, so the current tab is forgotten and later calls open new tabs (PreviewAutomationBroker.ts).

    Overlay registration already races hostDeadlineMs (the budget from #4685). Racing navigate against that deadline and giving readiness only the remainder matches that. #12407 is the same eviction mechanism for preview_wait_for; it is closed and does not cover this path. Resize is a lesser cousin: waitForRenderedViewport also starts a fresh full timeout, but it is bounded, unlike loadURL.

    4. Codex detach leaves the native thread and its MCP servers loaded (v2) — confirmed

    Codex is the only driver with supportsMultipleProviderThreadsPerSession: true. Detach on that path interrupts in-flight turns, then drops the app thread from attachedThreadIds and loadedProviderThreadKeyByThread. It never tells the runtime to unload the native thread, and ProviderAdapterV2SessionRuntime has no unload operation. The shared app-server stays up while any other thread is still attached (ProviderSessionManager.ts).

    CodexAdapterV2 sends thread/start and thread/resume and never thread/unsubscribe. The generated client already has that method; the response status is notLoaded | notSubscribed | unsubscribed, which is the unload signal. An optional unloadThread, implemented for Codex with thread/unsubscribe and called when detaching from a shared runtime, matches the leak. I did not reproduce the 80-process / 9 GB measurement or the 11-to-0 MCP check; those are yours.

    5. Failed turn stays "waiting" while background work is pending (v2) — confirmed

    shellRuntime forces runtime.status to "idle" whenever pendingBackgroundTasks is nonempty, including when the shell status is "failed" (packages/client-runtime/src/state/models.ts). resolveSidebarThreadStatus returns "waiting" for "idle" before it looks at "failed" (Sidebar.logic.ts). ThreadNotificationCoordinator only promotes a "ready" status when latestRun.status === "failed", so an idle/waiting failed thread never becomes the failure toast either.

    On main, the resolver checks a failed session before background liveness ("A failed session outranks lingering background liveness"). This branch inverted that. The existing test expects idle plus a leftover lastError to stay waiting, so the check should be "idle, and the latest run failed", not "any persisted lastError". resolveSidebarThreadStatus does not currently take latestRun; both the sidebar and the coordinator pass a shell that has it, so adding it to the pick fixes both. Idle with background work and a non-failed latest run should stay waiting.

  2. added
    acceptedfeature request accepted
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 24, 2026
  3. juliusmarminge commented on Sep 29, 2026

    @juliusmarminge
    Member

    Items 1, 4 and 5 have now been addressed

  4. letrandat commented on Oct 3, 2026

    @letrandat

    Adding a current observation of item 2: Lineage / “Previous agents” still shows “Done” while the same delegated worker is running follow-up work.

    Environment: T3 Nightly 0.0.46-nightly.20261003.2610 on macOS; installed bundle embeds commit 8ed276c. A native Claude coordinator delegates to a Codex worker.

    Observed sequence:

    1. The coordinator starts a worker thread. Its initial run completes, and the Lineage entry shows “Done”.
    2. The coordinator sends follow-up work to that same worker thread.
    3. During the follow-up, t3_thread_wait reports the target run as running, with no pending approval, and local output files continue updating.
    4. The coordinator's Lineage / “Previous agents” entry still shows “Done”.

    Reusing the existing worker thread is expected. The misleading part is the status: it appears to retain the initial run's completion while later work is active. The UI should reflect the latest run/current activity, or clearly distinguish historical completion from current activity.

    This was observed in an existing session. I have not established a separate clean reproduction or independently verified the source-level cause or a fix. No sensitive logs or screenshots are attached.

  5. rama-adi commented on Oct 3, 2026

    @rama-adi

    Confirmed item 2 with a rate-limit recovery trigger on T3 Desktop Nightly 0.0.46-nightly.20261003.2610, commit 8ed276c, macOS arm64, Claude Agent SDK.

    After Claude hit its rate limit, the user waited for reset and sent a new message to resume. A read-only database check showed:

    • A3 and G5: second runs running, parent subagent records still failed.
    • G4: second run completed, parent subagent record still failed.

    The parent's task queries returned the earlier rate-limit failure alongside hasPendingChildRuns: true.

    At this installed commit, packages/client-runtime/src/state/threadRelationships.ts:99 uses subagent.status for the edge, and apps/web/src/components/chat/ThreadRelationshipsControl.tsx:243 counts running agents from those records.

    Expected: Lineage reflects the child's current run while retaining the original failure as history.

    Related proposed fix: #12977.

    Diagnosed by GPT-6 via Codex.

  6. arhammahajan commented on Oct 3, 2026

    @arhammahajan
    Contributor

    I’m seeing another trigger for item 2: manually stopping and resuming a delegated subagent.

    Steps to reproduce

    1. Have a parent thread delegate work to a subagent.
    2. Check the parent thread’s Lineage: the subagent shows as running.
    3. Stop the subagent. Its Lineage status changes to Stopped.
    4. Resume the same subagent and confirm it is actively working again.
    5. Return to the parent thread’s Lineage.

    Expected behavior

    The subagent’s Lineage status returns to Running when it resumes.

    Actual behavior

    The subagent resumes work, but the parent thread’s Lineage continues to show Stopped.

    Environment

    Ubuntu 26.04

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions