Skip to content

App Crash on heavy thread #996

Description

@nassimna

my t3code app is crashing after i ran a code review thread on a monorepo project it just crash and sometimes it finishes the work, but i could not start any new thread on that workspace without deleting that thread, and the only solution to get the app working well again is to clear the .t3 folder

  • i have no attachment because i deleted the .t3 folder to fix the problem

Activity

  1. chuks-qua commented on Mar 12, 2026

    @chuks-qua
    Contributor

    Any rough idea how large the thread was? also it could be due to some rerendering issues too. @juliusmarminge recently applied fixes for those but without proper context Can't really do much to help.

    If you can provide a better estimate on the size of the thread that would be nice

  2. chuks-qua commented on Mar 12, 2026

    @chuks-qua
    Contributor

    related to #980
    Add this to ur desc please

  3. juliusmarminge commented on Mar 12, 2026

    @juliusmarminge
    Member

    i've pushed this pretty hard with long threads before so likely a regression. anything that helps debugging would be good

  4. nassimna commented on Mar 13, 2026

    @nassimna
    ContributorAuthor

    I tried to replicate the problem, but I couldn't after deleting my .t3 folder and reseting every thing but reading the #980 issue, I confirmed that had more than 20 worktrees between all my projects, and my .t3 was over 8 GB. Since there is no UI to manage worktrees in the current release, i dont really know how to manage that

  5. chuks-qua commented on Mar 13, 2026

    @chuks-qua
    Contributor

    I could not replicate the issue but I had claude investigate this deeply and eliminate the false positives too. here is what it found. the one that stood out the most is the snapshot which is likely the cause for #980. which is probably why @nassimna could access his workspace only after deleting the .t3 folder.


    Investigation: Root Cause Analysis

    I dug into the codebase to understand why this happens. Here's what I found.

    How worktrees work in T3 Code
    T3 Code creates real git worktrees (git worktree add) stored at ~/.t3/worktrees/{repoName}/{branch}`. The worktree path is tracked per-thread in the projection_threads SQLite table. Multiple threads can share one worktree.

    The core problem: worktrees accumulate and never get cleaned up

    Worktrees are only removed when a user explicitly deletes a thread from the sidebar. There is no automatic pruning, no count limit, no disk space check, and no startup cleanup. git worktree prune is never called anywhere in the codebase.

    Worse, there are several ways worktrees become permanently orphaned:

    • User declines the "Delete worktree too?" dialog when deleting a thread. The thread is already gone from the DB at that point, so the worktree can never be discovered again.
    • App crashes between thread deletion and worktree removal. The thread is marked deleted in the database (line 612 of Sidebar.tsx) before worktree removal is attempted (line 636). If the app crashes in that gap, the worktree is orphaned with no recovery path.
    • git worktree remove fails (permissions, timeout, disk full). A toast is shown, but the failed removal is never retried.
    • Project deletion doesn't cascade to threads. The project.deleted event only sets deletedAt on the project. Threads and their worktrees are left behind.
    • Database corruption or deletion orphans every worktree instantly, since the app never scans ~/.t3/worktrees/ on disk.

    Why accumulated worktrees cause freezes

    The getSnapshot() query (ProjectionSnapshotQuery.ts) loads ALL projects, threads, messages, activities, sessions, plans, and checkpoints in a single transaction with no WHERE deleted_at IS NULL filter. With many threads (including deleted ones that were never filtered out), this builds a large in-memory object graph on every snapshot refresh.

    The client requests a snapshot refresh on every domain event (throttled to 100ms). So during active agent sessions
    that emit many events, the snapshot query runs repeatedly with all that data.

    Additionally, listBranches() calls git worktree list --porcelain on every invocation, which slows down as worktrees accumulate.

    Each worktree with a setup script also gets its own node_modules (~1GB per worktree for a monorepo). This is the main source of the 8GB+ disk usage reported in #996.

    What's missing

    1. No automatic worktree cleanup (age-based, count-based, or on startup)
    2. No git worktree prune call anywhere
    3. No way to view or manage worktrees outside of deleting individual threads
    4. No cascading cleanup when projects are deleted
    5. No crash recovery for interrupted worktree removals
    6. Snapshot query loads deleted threads into memory unnecessarily

    @juliusmarminge i could whip up a pr (or multiple prs to match repo standards) attempting to fix these while you test when it's ready. what do you think?

  6. olafura commented on Jun 24, 2026

    @olafura

    @chuks-qua @nassimna Could you check out my branch to see if it fixes your problem

  7. EMOEMOJAI commented on Jul 20, 2026

    @EMOEMOJAI

    Reproduced with full diagnostics — backend OOM crash-loop, root cause confirmed

    Hit what looks like the terminal form of this bug on 0.0.29-nightly.20260718.841 (macOS 26.5.2, M4 Max): app shows "Failed to connect. Reconnecting… / cannot connect remote endpoint 127.0.0.1:3773" and crash-loops forever (25 macOS crash reports in one morning).

    Root cause: the backend child (apps/server/dist/bin.mjs) OOMs ~45s after every startup while loading thread history — consistent with the unfiltered snapshot load described in the investigation above:

    FATAL ERROR: MarkCompactCollector: young object promotion failed
    Allocation failed - JavaScript heap out of memory
    [30531] 44149 ms: Scavenge 3641.2 (3719.6) -> 3624.6 (3724.4) MB ...
    

    State at time of failure:

    • ~/.t3/userdata/state.sqlite: 5.1 GB; orchestration_events: 1,541,085 rows (~1.6 GB JSON payload); projection_thread_activities: 1.4 GB; 304 thread streams
    • Two single threads had 165k and 172k events each (~170 MB / ~150 MB payload) — long agent sessions
    • Backend RSS climbs to ~3.7 GB in ~45 s, then aborts at the default V8 heap cap; the desktop shell respawns it → infinite loop. In-memory expansion is roughly 5× the raw JSON payload.
    • NODE_OPTIONS=--max-old-space-size=... is not honored by the backend child (stripped), so there is no user-side escape hatch.

    Workaround that fixed it (no full .t3 wipe needed): with the app quit, delete old/huge thread streams from orchestration_events + matching rows in projection_* / orchestration_command_receipts / checkpoint_diff_blobs, then VACUUM. DB went 5.1 GB → 323 MB; backend now stable at ~300 MB RSS with no crashes.

    This is exactly the failure mode PR #3510 addresses — can confirm the pagination approach targets the right root cause. Happy to provide full crash reports or trace logs if useful.

  8. danieliser commented on Jul 22, 2026

    @danieliser

    Desktop backend OOM on large profiles: stale global replay and eager task hydration

    Summary

    A packaged macOS arm64 desktop build was crash-looping every 35–85 seconds against a 4.43 GiB state.sqlite containing approximately 1.24 million orchestration events, 432,000 projected thread activities, and multiple very large archived tasks.

    The backend child reached approximately 2,797,872 KiB RSS before aborting through node::OOMErrorHandler. The native stack passed through SQLite row extraction and V8 string creation.

    The database was not pruned. The stable solution was to decouple retained history from resident memory.

    Root cause

    The failure was not initial server startup or archived history merely existing on disk.

    Persisted client cursors were approximately 271,000–591,000 global events behind. The old subscribeThread resume path called:

    readEvents(afterSequence, Number.MAX_SAFE_INTEGER)

    It then filtered for the requested task only after SQLite had read and decoded every intervening global event.

    Tracing recorded 15,544 pages of:

    SELECT ...
    FROM orchestration_events
    WHERE sequence > ?
    ORDER BY sequence ASC
    LIMIT ?

    The sidebar also mounted hidden detail subscriptions for up to ten visible tasks. Each subscription could initiate another global catch-up, multiplying SQLite decoding and allocations.

    Opening a large task had another unbounded path: its complete activity history was materialized into the initial detail snapshot.

    What worked

    1. Integrated fix(server): paginate large thread history to stop the server running out of memory #3510:

      • Caps recent subscription replay at 1,000 global events.
      • Replaces stale-cursor replay with a fresh per-task snapshot.
      • Loads only the newest 500 activities initially.
      • Pages older activity through orchestration.getThreadActivities.
      • Adds on-demand older-history loading for web and mobile.
    2. Removed Electron sidebar detail prewarming:

      • Unopened tasks remain lightweight shell metadata.
      • Messages, activities, checkpoints, and task streams load when the task is viewed.
      • Existing task-detail state becomes eligible for disposal after five minutes without subscribers.
      • The current task and tasks with active observers remain subscribed normally.
    3. Fixed a correctness hole still present in fix(server): paginate large thread history to stop the server running out of memory #3510:

      • Its stale-cursor snapshot fallback omitted the requested synchronized marker.
      • The tested fix emits snapshot, then synchronized, then buffered live events, with a regression test.
    4. Added an 8 GiB packaged-backend heap default as defense in depth:

      • Explicit NODE_OPTIONS still wins.
      • Development and WSL environment plumbing are unchanged.
      • This is a safety rail, not the primary performance fix.
    5. Kept the earlier streaming improvements:

    Results

    Measurement Before Patched build
    Database ~4.43 GiB ~4.43 GiB, unchanged
    Backend RSS ~2,797,872 KiB before abort 295,056–437,216 KiB during soak
    Lifetime 35–85 seconds 15+ minutes and responsive
    Port 3773 Repeatedly disappeared Continuously listening
    Crash reports Roughly one per minute None during patched soak
    History deleted N/A None

    Observed backend RSS fell approximately 84–89% without shrinking the database.

    No event, activity, checkpoint, message, or archived-task rows were deleted. No VACUUM or log rotation was performed. At the end of verification the live database was 4,755,988,480 bytes and the existing log directory was still 6.6 GiB.

    Verification

    • 146 focused tests passed across desktop configuration, server projection/paging, stale-cursor recovery, shared lazy-history state, and web behavior.
    • The existing 22 streaming projection tests also pass.
    • Targeted type checks passed for contracts, client-runtime, web, desktop, mobile, and server.
    • Targeted lint and formatting checks passed.
    • An arm64 DMG was rebuilt and installed into /Applications.
    • With the parent environment unset, the backend child received NODE_OPTIONS=--max-old-space-size=8192.
    • Live use confirmed the packaged application is substantially more responsive.

    Upstream status and remaining work

    This is not fully solved on main.

    The important result for this issue is that deleting .t3 or pruning old chats is a workaround, not a required fix. The same large database became stable once hydration and replay were bounded.

  9. saphid commented on Jul 29, 2026

    @saphid
    Contributor

    Reproduction on 0.0.30-nightly.20260729.938: authenticated full snapshot deterministically OOMs the backend

    I captured a direct request-level reproduction on macOS arm64 that narrows this to the global snapshot hydration path.

    Trigger

    An authenticated curl/8.7.1 request called:

    GET /api/orchestration/snapshot
    

    against a valid retained profile. The trace recorded:

    state.sqlite                         1,589,829,632 bytes
    projection_thread_activities              90,145 rows
    activity payload/summary/kind JSON      651,155,472 bytes
    orchestration_events                     124,944 rows
    
    largest SQL execute                   2,331.504 ms
    SQL transaction                       3,443.996 ms
    environment.orchestration.snapshot    4,107.530 ms
    HTTP request                           5,127.946 ms (400)
    

    Eighteen seconds after the failed response, the backend child aborted:

    Mark-Compact (reduce) 3325.7 (3367.8) -> 3325.7 (3367.3) MB
    FATAL ERROR: CALL_AND_RETRY_LAST Allocation failed - JavaScript heap out of memory
    

    The native stack passes through:

    node::sqlite::StatementExecutionHelper::ColumnToValue
    node::sqlite::ExtractRowValues
    node::sqlite::StatementExecutionHelper::All
    

    The desktop supervisor restarted the backend one second later. The retained child log contains six matching heap-OOM incidents between July 26 and July 29.

    Why this is useful

    This request exercised ProjectionSnapshotQuery.getSnapshot(), which eagerly reads all projects, threads, messages, plans, activities, sessions, checkpoints, latest turns, and projection state into one object graph. About 651 MB of activity projection reached the roughly 3.3 GB V8 heap ceiling while SQLite rows were being converted.

    This is not a slow-link case and not an inferred renderer leak: the failing process was the backend child, the route and trace are known, and one authenticated GET was sufficient. The lightweight /api/orchestration/shell route remained fast after the supervisor restart.

    The profile and event history remain intact. I did not trim projection rows, delete events, or wipe .t3.

    This also triggered the existing startup-reconciliation problem in #4584: active turns became ghost running sessions after the backend respawned. I am adding that evidence there separately.

    Related: #4595 and #3510.

  10. danieliser commented on Jul 29, 2026

    @danieliser

    @saphid did you try the solutions I posted? I've been crash free since I posted that. Some chats running for days.

  11. saphid commented on Jul 29, 2026

    @saphid
    Contributor

    @saphid did you try the solutions I posted? I've been crash free since I posted that. Some chats running for days.

    I had not, doing now.

  12. danieliser commented on Aug 4, 2026

    @danieliser

    @saphid Any update?

  13. t3-code commented on Aug 4, 2026

    @t3-code
    Contributor

    Closing as fixed by merged PR #5147, which bounded event replay and removed full snapshot hydration from the affected paths: #5147

  14. ahansen1234 commented on Aug 4, 2026

    @ahansen1234

    I am also seeing this bug. This ticket is closed but I am still seeing the issue

    Summary

    T3 Code 0.0.31 repeatedly crashes its backend process when loading a thread with a large accumulated tool-activity history.

    This is not a renderer or network failure. The embedded Node backend reaches V8's heap limit while materializing every projection_thread_activities row for one thread using SQLite StatementSync.all().

    The resulting UI message ("did not respond during connection setup") occurs because the backend aborts and its WebSocket disappears.

    Area

    apps/server

    Impact

    Major degradation or frequent failure. Once the affected thread is restored or opened, the backend enters an OOM crash loop and T3 Code becomes unusable.

    Environment

    • T3 Code: 0.0.31
    • Codex CLI: 0.145.0
    • macOS: 15.6.1 (24G90)
    • Architecture: Apple Silicon / arm64
    • RAM: 48 GiB
    • SQLite database: ~/.t3/userdata/state.sqlite

    Steps to reproduce

    1. Accumulate a large amount of tool activity in a single thread, especially large command output, file reads, diffs, or other tool.completed payloads.
    2. Quit and reopen T3 Code, or navigate back to the large thread.
    3. T3 loads the complete thread detail, including every activity payload.
    4. Backend memory grows to approximately 3.4–3.6 GiB.
    5. V8 aborts the backend with JavaScript heap out of memory.
    6. The UI reports a generic connection/setup failure.

    In my case, the affected thread was also still marked running, so it continued accumulating activity until T3 Code was fully stopped.

    Actual behavior

    The server trace shows this unbounded query:

    SELECT
      activity_id AS "activityId",
      thread_id AS "threadId",
      turn_id AS "turnId",
      tone,
      kind,
      summary,
      payload_json AS "payload",
      sequence,
      created_at AS "createdAt"
    FROM projection_thread_activities
    WHERE thread_id = ?
    ORDER BY
      CASE WHEN sequence IS NULL THEN 0 ELSE 1 END ASC,
      sequence ASC,
      created_at ASC,
      activity_id ASC
    
    There is no `LIMIT` or pagination.
    
    The native crash stack contains:
    
    node::sqlite::StatementSync::All(...)
    node::sqlite::StatementExecutionHelper::All(...)
    node::sqlite::ExtractRowValues(...)
    node::sqlite::StatementExecutionHelper::ColumnToValue(...)
    v8::String::NewFromUtf8(...)
    node::OOMErrorHandler(...)
    abort()
    
    The backend log reports:
    
    FATAL ERROR: MarkCompactCollector: young object promotion failed
    Allocation failed - JavaScript heap out of memory
    
    The failed backend launched at approximately `10:52:21` and aborted at `10:55:14`, about 174 seconds later.
    
    ### Database evidence
    
    The database is large but valid:
    
    state.sqlite size:             1,747,128,320 bytes
    PRAGMA quick_check:            ok
    
    orchestration_events:          797.2 MiB
    projection_thread_activities:  710.4 MiB
    
    One thread accounts for:
    
    Projected activities:          60,522 rows
    Projected payload size:        321.6 MiB
    Largest projected payload:     1.2 MiB
    
    thread.activity-appended:      60,568 events
    Event payload size:            337.0 MiB
    
    tool.completed activities:     16,581 rows
    tool.completed payload size:   310.9 MiB
    
    Projected messages:            2,525
    Total message text:            0.84 MiB
    
    This means the actual conversation text is small. Almost all of the size comes from tool-activity payloads stored in both the projection and event log.
    
    `StatementSync.all()` first materializes hundreds of megabytes of strings and row objects. JSON decoding/schema validation and the thread snapshot then multiply the retained heap enough to exceed V8's limit.
    
    ### Expected behavior
    
    Loading a large thread should not materialize its complete activity history in one operation.
    
    The server should:
    
    - Return only the newest bounded activity window, with a `hasMoreActivities` indicator.
    - Paginate or lazy-load older activity.
    - Avoid selecting/decoding large `payload_json` values until requested.
    - Preserve unresolved approvals and user-input requests even when truncating history.
    - Avoid prewarming complete details for hidden or archived threads.
    - Consider externalizing or otherwise bounding unusually large tool results.
    
    Messages and checkpoints can remain complete because they are not the dominant data here.
    
    ### Related work
    
    PR #5147 fixed two other OOM paths: stale-cursor global event replay and full-database snapshot hydration:
    
    https://github.com/pingdotgg/t3code/pull/5147
    
    This report is a different remaining path: direct thread-detail hydration through `listThreadActivityRowsByThread` / `getThreadDetailById`.
    
    A previous proposed fix capped initial history at the newest 500 activities, but that PR was closed without merging:
    
    https://github.com/pingdotgg/t3code/pull/4948
    
    I also checked `v0.0.32-nightly.20260804.998`; its per-thread activity query still uses `SqlSchema.findAll` without a `LIMIT`, so updating to that nightly does not resolve this particular crash.
    
    ### Workaround
    
    Increasing `--max-old-space-size` may allow the thread to load temporarily, but does not fix the unbounded query.
    
    The reliable local workaround is to:
    
    1. Fully stop T3 Code and its backend/provider children.
    2. Make a verified SQLite backup.
    3. Remove the affected thread's projected activity rows and corresponding `thread.activity-appended` events.
    4. Stop its stale running session.
    5. Checkpoint and `VACUUM` the database.
    
    This preserves the thread's messages and filesystem/worktrees but removes its tool-activity timeline.
    
    ### Diagnostics available
    
    I can provide:
    
    - The macOS `.ips` crash report.
    - A sanitized `server-child.log` excerpt containing the V8 OOM.
    - A sanitized trace containing the exact SQL query.
    - Aggregate SQLite counts and payload lengths.
    
    I cannot provide the complete SQLite database because its tool payloads may contain private source code, command output, paths, or credentials.
    
    Attach the `.ips` crash report if possible, but inspect and sanitize logs before uploading. Do not attach `state.sqlite`; it contains conversation and tool-output data.
  15. danieliser commented on Aug 4, 2026

    @danieliser

    This fix is still relevant for me on the nightly from yesterday. Either the new bounded fix isn't the whole picture, or it hasn't been pushed out in a nightly?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.🚧 In Progress

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions