Repository navigation
Codex diff snapshots can accumulate in unbounded ingestion queues #12883
Description
Activity
Triage
Confirmed on current
main(1de563c14, after the reported nightly7445aa733ada). This is a real unused-payload retention path. It is not fixed by #12305, and it is not a duplicate of #12758 or #12884.#12305 only bounds diagnostic provider log records. Native
turn/diff/updatedframes are now skipped entirely (transientNativeMethods). Canonicalturn.diff.updatedis still persisted, but #12305 clamps the serialized line. The runtime objects that feed the queues are unchanged.What the code does
CodexAdaptermaps a nativeturn/diff/updatedonto a canonical event that keeps the same string twice:raw.payload.diffviaruntimeEventBasepayload.unifiedDiffin theturn/diff/updatedmapper
(CodexAdapter.ts~975–993, ~1664–1676)
- That full event then sits in these unbounded buffers:
CodexSessionRuntime:Queue.unbounded<ProviderEvent>()CodexAdapter:Queue.unbounded<ProviderRuntimeEvent>()ProviderService:PubSub.unbounded<ProviderRuntimeEvent>()ProviderRuntimeIngestiondiffWorkerand, after repo detection, the lifecycle worker (makeDrainableWorker→TxQueue.unbounded)
recordProviderDiffnever readsunifiedDifforraw.payload.diff. It dispatches a placeholder checkpoint:status: "missing",files: []. After the first placeholder for a turn, later snapshots are dropped (hasCheckpointForTurn).- fix(server): keep VCS waits from blocking turn completion #11970 moved
checkpointStore.isGitRepositoryonto the dedicateddiffWorkerso a hung git no longer stalls turn completion. That also means the worker can accept snapshots faster than it can drain them, and each queued item still holds the full strings. - Clients never read
unifiedDiff. The field exists only on the contracts schema (TurnDiffUpdatedPayload) and in server fixtures/tests.
CheckpointReactoralso subscribes toProviderService.streamEventsbut immediately dropsturn.diff.updated. PubSub still delivers the full object to that subscriber.Only Codex emits this event type.
What this is not
The report is careful, and the code agrees: this is a reachable unbounded path, not a proven crash cause. Provider logs do not record queue depth. The ~49 million-character diffs cited here are the durable Codex file-change payloads from #12758, which shows this provider can emit huge diff strings on the affected install. They are not measured
turn/diff/updatedqueue entries.Related, not duplicates:
- Oversized Codex diff event exhausts desktop backend heap #10924 / fix(server): bound provider event log records before serialization #12305 — logger bound only. The Oversized Codex diff event exhausts desktop backend heap #10924 triage already noted orchestration does not store
unifiedDiffand asked to stop the dual-field copy. fix(server): bound provider event log records before serialization #12305 did not do that. - fix(server): stop an oversized provider diff from exhausting the backend heap #10965 / fix(server): stop an oversized provider diff from exhausting the backend heap #10968 / fix(server): stop an oversized provider diff from exhausting the backend heap #11125 (closed) — would have replaced
raw.payload.diffwith a marker and still kept the fullunifiedDiffon the runtime event. - fix(server): keep VCS waits from blocking turn completion #11970 — isolated the VCS wait; that is the accumulate-while-git-is-slow path.
- Codex file-change diffs are duplicated in SQLite and exhaust backend heap #12758 — durable SQLite copies of
item.started/item.completedfile-changechanges[].diff. - Codex app-server input has no per-message byte limit #12884 — missing app-server per-message byte limit before parse.
- Background service exhausts V8 heap and aborts after ~1-2 days of continuous uptime under sustained session load #12584 / [Bug]: Desktop backend still SIGABRTs with V8 OOM after 9–18h on 0.0.35 (hydration bounds already present) #8648 — broader long-uptime / large-profile heap exhaustion.
- [Bug]: app-server stdin reader is quadratic in line length — OOM crash on large tool payloads #5389 — quadratic app-server reader. Fixed.
- Checkpoint diffs never resolve: ProviderRuntimeIngestion placeholder preempts CheckpointReactor git capture #585 — placeholder checkpoint ordering, not memory bounds.
Next step
Accept. Slim the event to the fields
recordProviderDiffactually uses (event / thread / turn identity and time) before the first queue, or at least beforediffWorker.enqueue. Do not keepraw.payload.difforpayload.unifiedDiffon queued work. Coalesce repeated snapshots for the same thread and turn (KeyedCoalescingWorkeralready exists). A later snapshot after the first placeholder is already a no-op.Add a regression that a generated large
params.diffnever remains on queued ingestion work, and that many updates for one thread+turn stay bounded.Restarting the backend clears the live queues. It does not close the path.
Not closing.
- addedacceptedfeature request acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 21, 2026 Note
🤖 GPT-6.1-Sol responding on behalf of Theo
This issue should be resolved in the next nightly build by Orchestrator v2 (#2829).
Please try that nightly. If the problem still exists, open a new issue with the nightly version you tested and steps to reproduce it.
What happened
Large Codex
turn/diff/updatedsnapshots remain attached to runtime events after provider event logging has bounded its output. T3 then sends each full event through unbounded queues before it records a small placeholder checkpoint.The ingestion code never reads
payload.unifiedDiff. It uses only event identity, thread and turn identity, and time. Keeping the full diff in these queues adds memory use without changing behavior.This source audit followed repeated desktop backend V8 out-of-memory failures. The logs do not record queue depth, so this report does not claim that queue growth caused those failures. It reports a reachable, byte-unbounded path found during that investigation.
Diagnosis
CodexAdaptermaps a nativeturn/diff/updatednotification to a canonical event that retains the diff in two fields:raw.payload.diffpayload.unifiedDiffThe event then passes through these unbounded buffers:
ProviderService's runtime event pub-sub buffer.ProviderRuntimeIngestion's diff worker.makeDrainableWorkerusesTxQueue.unbounded. The diff worker can wait for repository and VCS checks, so its consumer can run more slowly than the provider event source.recordProviderDiffdoes not readunifiedDiff. It creates a placeholder checkpoint withstatus: "missing"andfiles: []. The full string can be removed before the first queue, or at least beforediffWorker.enqueue.Provider logs from the affected install contain 3,017 canonical diff notifications after the logger fix, with a peak of 57 notifications in one minute. The logger records are bounded, so they do not reveal the original diff sizes. Separate read-only database checks found individual decoded Codex file-change diffs near 49 million characters, which confirms that this provider can produce large diff strings on this install.
Expected behavior: T3 should not retain unused full diff strings in unbounded queues. It should reduce the event to the fields required for checkpoint scheduling, coalesce repeated snapshots for the same thread and turn, or apply bounded backpressure.
Steps to reproduce
Source-level reproduction:
turn/diff/updatednotification with a generated large string inparams.diff.CodexAdaptermap it to a canonical runtime event.checkpointStore.isGitRepositoryor otherwise pause the diff worker.ProviderService.streamEvents.diffWorkeraccepts all events and retains their fullraw.payload.diffandpayload.unifiedDiffstrings.A regression test can use a small configured queue or generated strings. It should assert that queued diff work does not retain
unifiedDiffand that repeated updates for one thread and turn stay bounded.Version
0.0.43-nightly.20260920.2005, commit7445aa733ada.Environment
Evidence
Related issues
No open issue or pull request found in the upstream search covers unused full diffs in the ingestion queues.
Fix applied or workaround
No source patch or state change was made. Restarting the backend clears queued events but does not prevent the path from recurring.
Filed by
Codex, GPT-5, via
t3 triage