Skip to content

Process externally queued Codex messages without opening the T3 thread #12448

Description

@praveenperera

What happened

An external process added a message to an existing T3-managed Codex thread with codex queue. Codex stored the message, but T3 Code did not process it until the user opened the thread in the T3 Code mobile app.

The thread was not settled. The callback waited for 5 hours, 48 minutes, and 52 seconds even though T3 Serve remained available.

A queued Codex message should activate its T3-managed thread without requiring a browser or mobile client to open it. If the thread is settled, T3 should unsettle it when processing starts. If the thread is already active, T3 should still restore its missing provider session and process the queue.

Diagnosis

codex queue stores the message in Codex's durable external queue. The Codex queue watcher dispatches messages only for threads that are loaded in an app-server process.

T3 Code had no loaded Codex provider session for this thread. It did not load or resume the session when the external message was queued. Opening the thread caused T3 to restore the provider session, after which Codex consumed the queued message.

This is a provider lifecycle wake gap. It is not message loss, and it does not depend on the T3 thread being settled.

T3 Code needs either automatic detection of queued Codex messages or an authenticated wake operation for external callers. A wake should:

  1. Resolve the T3 thread from its Codex thread ID.
  2. Load or reconnect the Codex provider session.
  3. Resume the Codex thread.
  4. Let Codex consume the message that is already in its queue.
  5. Show the thread as running while processing occurs.
  6. Unsettle the thread if it is settled.
  7. Remain safe when called more than once.

The wake operation must not submit the callback text again because that would risk a duplicate turn.

Steps to reproduce

  1. Start T3 Serve and create or resume a Codex thread through T3 Code.
  2. Leave the thread so that its Codex provider session is no longer loaded.
  3. From another process, run codex queue --thread <thread-id> --message <text>.
  4. Do not open the thread in a browser or mobile client.
  5. Observe that the queued message is not processed.
  6. Open the thread in T3 Code.
  7. Observe that T3 loads the provider session and the queued message starts processing.

Version

T3 Code 0.0.43-nightly.20260917.1866

Environment

Linux x86_64, T3 Serve, T3 Code mobile client, Codex CLI 0.154.0

Evidence

External work started:       10:46:05 PM CDT (UTC−05:00)
External work completed:     12:42:06 AM CDT (UTC−05:00)
Callback queued:             12:42:22 AM CDT (UTC−05:00)
Callback processed by T3:     6:31:14 AM CDT (UTC−05:00)
Queue-to-processing delay:    5:48:52

Processing started only after the thread was opened in the T3 Code mobile app.

Related issues

#5123 requests an atomic guarded operation that submits a new turn to an idle thread. This report is different: the message already exists in the Codex queue, the T3 thread was not settled, and T3 must restore the provider session so Codex can consume the queued message.

Fix applied or workaround

Opening the thread causes T3 Code to load the provider session and process the queued message. This requires user action and does not support unattended callbacks.

Filed by

Me & Codex 5.6 sol

Activity

  1. juliusmarminge commented on Sep 18, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed as a provider-lifecycle wake gap, not message loss and not a settle-state bug. Not a duplicate of #5123 or #10183.

    codex queue writes Codex’s durable external queue. Codex only dispatches those items for threads that are loaded or resumed in a running app-server (openai/codex#39034, openai/codex#39092). T3 starts codex app-server per thread and calls thread/resume only when it creates a provider session. Idle sessions are reaped after 30 minutes (ProviderSessionReaper). There is no T3 path that notices the external queue or restores that session without a user turn.

    Opening the thread in a client only subscribeThreads (read-only). Session start is a side effect of thread.turn.start, runtime-mode change, compaction, or server-restart continuation. A view-only open should not load Codex; the observed workaround may have included a later send.

    Why this is not #5123 or #10183

    Where this lives

    • apps/server/src/provider/Layers/CodexSessionRuntime.ts — thread/resume on session start
    • apps/server/src/orchestration/Layers/ProviderCommandReactor.ts — ensureSessionForThread
    • apps/server/src/provider/Layers/ProviderSessionReaper.ts — 30-minute idle stop
    • apps/server/src/orchestration/http.ts — authenticated dispatch only; no restore-only wake
    • No lookup of a T3 thread by Codex providerThreadId
    • Vendored app-server schema has no thread/queue/*

    Suggested direction

    Enhancement, medium. Smallest fix: an authenticated restore-only wake (not thread.turn.start):

    1. Resolve the T3 thread from the Codex thread id
    2. startSession / ensureSessionForThread with the persisted resume cursor
    3. Let Codex consume the already-queued item
    4. Unsettle if settled; show running while it processes
    5. Stay idempotent
    6. Do not submit the callback text again

    ProviderService.startSession already reapplies the binding resume cursor. Treat auto-watching ~/.codex/queue_*.sqlite as a later option; it couples T3 to Codex’s queue schema.

    Surfaces

    Server/Codex adapter only. Web, desktop, and mobile would all see the thread run once the session is restored. No provider-specific UI. Docs should mention that codex queue against a T3-managed thread needs this wake (or a later auto-watch) after the session has been reaped.

  2. added
    enhancementRequested improvement or new capability.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 18, 2026
  3. praveenperera commented on Sep 27, 2026

    @praveenperera
    Author

    This comment was written by Claude Opus 5.5 through T3 Code on behalf of Praveen Perera.

    Workaround in homebased

    I fixed this on my side in homebased, the daemon that sends callbacks to my agent threads. It uses T3 Code's undocumented local HTTP API.

    What it does

    When a callback targets a Codex thread that T3 owns, homebased no longer uses codex queue. It sends the callback to T3 as a normal turn:

    1. Read ~/.t3/userdata/server-runtime.json to find the running server (origin and pid).
    2. Find the T3 thread in state.sqlite. The query matches provider_session_runtime rows where provider_name = 'codex' and json_extract(resume_cursor_json, '$.threadId') equals the Codex thread ID. It joins projection_threads to skip deleted threads.
    3. Issue a short-lived bearer token with t3 auth session issue --json --ttl 5m --label homebased. homebased runs the server's own t3 binary, which it finds through /proc/<pid>/exe or ps, because the macOS desktop app does not put t3 on PATH.
    4. Call GET /api/orchestration/threads/{threadId}?turnLimit=1 to read runtimeMode and interactionMode.
    5. Call POST /api/orchestration/dispatch with a thread.turn.start command that contains the callback text. commandId and messageId are UUIDs that come from a hash of the thread ID and the text, so a retry does not make a second turn.
    6. Revoke the token with t3 auth session revoke <sessionId>.

    thread.turn.start makes T3 restore the provider session with the saved resume cursor. The turn then runs and shows in every client. If T3 cannot take the turn, homebased falls back to codex queue and sends one push notification that tells me to open the thread.

    The same path also wakes Claude Code sessions after T3 stops their idle process.

    Why I did it this way

    • The callback text goes to T3 or to codex queue, never to both. This prevents the duplicate turn that the triage warns about. The result is close to the restore-only wake that the triage suggests, but it is not identical.
    • Unattended callbacks waited for hours until I opened the thread. I needed a fix now, not after an upstream change.
    • thread.turn.start is the path that the T3 UI already uses, so the turn shows the same as a message that I typed.

    These endpoints are not a public contract, so homebased pins the behavior it observed (T3 Code commit 95030dc67, stable 0.0.42, nightly 0.0.43). homebased t3 check probes the contract without starting a real turn. The daemon runs this check again when the T3 server restarts and sends a push notification if the API changed. I would still like a supported, authenticated wake or turn endpoint for external callers so that this does not depend on internal details.

    Commits

    • f68e39a: Add a T3 Code client and API check
    • dafd442: Wake stopped Claude sessions through T3 Code
    • e77cce3: Bound T3 CLI calls with a deadline
    • b67244b: Deliver T3-owned Codex callbacks through T3 (the fix for this issue)
    • a01c28d: Find T3 when it runs as a service or on macOS

    The client and the pinned contract are in src/t3.rs.

  4. mauricekleine commented on Oct 7, 2026

    @mauricekleine

    We hit the same gap with Claude Code threads, so this isn't only a Codex problem.

    Our setup: a small local relay wakes coding agents when someone mentions them in a team chat. For Claude Code, it writes to the session's messaging socket, so the process has to be running. T3 releases an idle Claude session after 30 minutes. After that, nothing outside T3 can start the next turn.

    Last night on 0.0.46-nightly.20261003.2610 (macOS): the session went idle at 22:12 UTC and T3 released it at 22:42. A message for it came in at 22:47, then sat there until I typed in the thread at 22:52.

    Two things we tried and dropped:

    What would fix it for us, for any provider:

    1. A local way to wake one thread: a same-user socket, a t3 thread wake <threadId> command, or a token scoped to orchestration:operate on that thread only.
    2. A lookup from the provider's id to the T3 thread. For Claude, that's the --resume session id.
    3. A send with an idempotency key (clientRequestId). Restore-only, as the triage suggests, covers Codex. A restored Claude session still needs someone to start the turn, so a send is the simpler contract for both.

    Happy to build against that instead of state.sqlite and the dispatch payload.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedenhancementRequested improvement or new capability.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions