Skip to content

[Bug]: Checkpoint capture on large monorepos retries a guaranteed 30s git-add timeout every turn — permanent CPU burn + tmp_pack disk litter #3646

Description

@mikeohuo

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Open a workspace inside a large monorepo (in my case a Unity project — large pack file, very high file count; workspace cwd is a subdirectory of the repo root). The larger the repo, the worse this gets.
  2. Run agent threads normally. Every completed turn triggers CheckpointReactor → GitVcsDriver.checkpoints.captureCheckpoint, which runs git read-tree HEAD + git add -A -- . against a temp GIT_INDEX_FILE.
  3. Watch Task Manager / .git/objects/pack/.

Expected behavior

Checkpoint capture should either succeed, or fail once and back off / disable itself for the workspace. A capture that cannot ever succeed should not be retried indefinitely, and killed git processes should not leave permanent garbage in .git.

Actual behavior

VcsProcess has DEFAULT_TIMEOUT_MS = 30_000 and captureCheckpoint passes no timeoutMs override. On a large monorepo a full-tree git add -A takes well over 30 seconds, so the process is killed at the timeout every single time. The checkpoint is recorded as missing (every thread.turn-diff-completed event in my state.sqlite has status="missing", across every thread; no refs/t3/checkpoints/* ref has ever been created), and the next turn simply retries. Net effect:

  • One CPU core pegged near-continuously by git for as long as any thread is in use — capture can never succeed on this repo, so the cost buys nothing.
  • Each kill that lands mid-stream of a blob larger than core.bigFileThreshold (files that size are common in Unity repos) leaves an orphaned tmp_pack_* in .git/objects/pack/. These accumulate over time — git count-objects -v flags them as garbage, along with a large number of unreferenced loose objects from the partial adds.

Related issues — this seems to be the connecting root cause:

Suggested direction (any one of these would resolve the pathology): a per-project or global setting to disable checkpointing; skip-with-backoff after N consecutive capture failures instead of retrying forever; an adaptive/longer timeout; and cleanup of tmp_pack_* left by killed VCS processes.

Impact

Major degradation or frequent failure

Version or commit

0.0.29-nightly.20260701.697 (desktop, Windows installer)

Environment

Windows 11 Pro (10.0.26200), T3 Code Nightly desktop, git 2.54.0.windows.1, provider claudeAgent

Logs or stack traces

Observed process loop (spawned by the t3code server process, back-to-back):
  git.EXE -C <repo>/<workspace-subdir> add -A -- .    (killed at ~30s)
  git.EXE -C <repo>/<workspace-subdir> add -A -- .    (respawned on next turn, killed at ~30s)
  ... repeats indefinitely; ~100% of one core per run

state.sqlite orchestration_events: every thread.turn-diff-completed has status="missing"

git count-objects -v:
  warning: garbage found: .git/objects/pack/tmp_pack_* (many, accumulating)
  (plus a large count of unreferenced loose objects)

Workaround

None found — there's no setting that disables checkpointing. Locally I'm considering patching captureCheckpoint in the bundled bin.mjs to early-exit for this repo, but a nightly update overwrites it.

Activity

  1. AlexanderDotH commented on Aug 11, 2026

    @AlexanderDotH

    Additional Linux desktop reproduction:

    • I confirmed that the T3 Code server process directly spawns git -C add -A -- .; the application, not the target project, owns the Git child.
    • The workspace had one unignored ELF core dump, 32 GiB apparent size and 4.3 GiB allocated. That is enough to make every automatic checkpoint attempt expensive.
    • While observing the failure, Git garbage grew from 132 to 142 tmp_pack_* files. The new files were about 790-803 MiB each, with final mtimes spaced almost exactly 30 seconds apart.
    • Current source still has the same chain: CheckpointReactor captures a pre-turn baseline from message-sent, turn-start-requested, and turn.started events; GitVcsDriver runs add -A -- .; VcsProcess supplies the 30,000 ms default timeout; and capture cleanup removes only the temporary index. I found no source-side cleanup for stale tmp_pack_* files.

    I did not capture the server timeout log itself, but the live process ancestry, exact command, source path, and 30-second cadence make the timeout/retry explanation very strong. This reproduces outside a large monorepo too: a single large, unignored crash dump is sufficient.

    An adaptive timeout alone would still put crash dumps into hidden checkpoint history. A safer resolution would make capture failure back off for an unchanged workspace, expose an actionable reason, and clean only temp packs created by a fully terminated failed capture.

  2. mikeohuo commented on Sep 7, 2026

    @mikeohuo
    Author

    Follow-up reproduction on an official nightly

    This still reproduces when the Git root, T3 project directory, and backend working directory are all the same directory. The agent completes its work, but its conversation turn has no usable checkpoint diff.

    Environment and scope

    • Windows; a large repository on a secondary drive.
    • Official installed nightly: 0.0.39-nightly.20260906.1316.
    • The unmodified installed server bundle ran with an isolated T3 profile and its working directory set to the repository root.
    • The T3 project used the same repository root, with .git directly inside it.
    • The checkout already contained unrelated uncommitted changes and a linked worktree directory. This was not a clean, minimal reproduction.

    Observed result

    A claudeAgent turn completed successfully and created the requested small text file. Existing unrelated worktree changes were also present.

    Native vcs.refreshStatus, review.getDiffPreview, and file-content requests returned data. The proof file's ordinary diff contained the exact expected text. These were backend/provider checks; I did not visually verify the rendered UI.

    For this one turn, the server logged three separate checkpoint timeouts. Each reported the following error, with the repository path redacted:

    VcsProcessTimeoutError: VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repository-root>) after 30000ms
    

    The next warning was:

    checkpoint capture missing pre-turn baseline
    

    No checkpoint ref was created for the test thread. Its checkpoint list remained empty. orchestration.getTurnDiff failed with CheckpointTurnRangeUnavailableError, including:

    Turn diff range exceeds current turn count: requested 1, current 0.
    

    Four empty temporary checkpoint-index .lock files remained afterward. This run did not measure CPU utilization or observe new tmp_pack_* files; those remain separate observations from the earlier reports.

    What this separates

    The checkpoint timeout persists with all three directories aligned at the Git root. The separate review-directory validation issue was not blocking these requests.

    The practical failure here is missing per-turn conversation diffs after successful agent edits. Ordinary Git diff access does not establish that checkpoint-based diffs work.

    Proposed next steps

    Please prioritize successful checkpoint capture on large dirty repositories, with bounded retries and cleanup as supporting fixes.

    Several existing PRs address different parts of this:

    Useful acceptance coverage:

    1. Test a large dirty repository with both root and nested project directories.
    2. Cover staged, unstaged, deleted, and untracked files, including a newly staged file subsequently deleted from disk.
    3. Confirm that two consecutive turns create checkpoints and return the exact changes for each turn.
    4. Confirm that checkpoint operations preserve the user's real Git index and unrelated files.
    5. On failure, bound repeated attempts, clean only the failed capture's own temporary storage, and show a useful failure reason.

    I can retest an official nightly containing a fix against this repository.

    Audit attribution

    This audit was performed by Astra and Claude Fable 5.1 agents. Astra performed the source inspection and native endpoint checks. Fable 5.1 executed the live agent smoke test and independently reviewed the findings.

  3. axelcursive commented on Sep 11, 2026

    @axelcursive

    Additional Linux/NFS reproduction on T3 Code 0.0.40:

    Environment

    • T3 server running on a remote Linux machine
    • client connected remotely
    • repository stored on an NFSv4 volume
    • approximately 28,000 tracked files
    • .git directory approximately 2.6 GB

    Observed behavior

    An otherwise successful agent turn ended with:

    VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<workspace>) after 30000ms

    The response and working-tree changes remained available. The failure therefore appeared non-blocking, but the checkpoint for that turn was unavailable and the associated diff/revert may be incomplete.

    There was no Git lock, and the repository remained usable. git status --porcelain completed in approximately 2.3 seconds.

    Timings

    I reproduced the checkpoint sequence with a temporary GIT_INDEX_FILE and timed its main operations independently:

    • git read-tree HEAD: approximately 0.8 seconds
    • git add -A -- .: approximately 21–24 seconds without heavy contention
    • git write-tree: approximately 0.3 seconds

    At the time of the original failure, several I/O-heavy transfers were using the same shared NFS volume. The git add -A step was already close enough to the fixed 30-second deadline that this additional contention appears sufficient to explain the timeout.

    Temporary checkpoint index files were cleaned up correctly in this reproduction, and I found no indication of repository corruption.

    Limited comparison

    While investigating, I briefly inspected the public openai/codex repository. Its per-turn diff tracker appears to retain exact deltas from recorded apply_patch mutations without rereading the entire workspace. This is only a limited observation, not a suggested implementation: I did not study all of its editing and restoration paths, its desktop client is not fully represented in the public repository, and arbitrary shell edits involve different tradeoffs.

    I am mentioning it only as evidence that change-tracking designs can have different performance characteristics. The appropriate checkpoint semantics for T3 are best left to the maintainers.

  4. stychu commented on Sep 14, 2026

    @stychu

    Additional macOS reproduction on Apple Silicon. Notable because the repo is ordinary.

    The existing reports all involve something extreme: a Unity monorepo, a 32 GiB unignored core dump, an NFSv4 volume. This one is a plain local-APFS checkout of pingdotgg/t3code itself, and it still fails.

    Environment

    • macOS, Darwin 25.6.0, arm64, local APFS SSD
    • git version 2.50.1 (Apple Git-155)
    • Repository: pingdotgg/t3code @ main, 21,789 tracked files, .git 124 MB
    • Workspace cwd is the repo root; no linked worktrees, no hooks, no core.fsmonitor / untrackedCache / preloadindex set
    • git status --porcelain returns in well under a second; no stale index.lock

    Timings

    Reproduced the capture sequence against a temporary GIT_INDEX_FILE, warm cache, machine otherwise idle:

    step real
    git read-tree HEAD 0.11s
    git add -A -- . 6.7–8.0s
    git write-tree 0.08s

    Same shape as the NFS report above (0.8s / 21–24s / 0.3s), just proportionally smaller.

    Isolating the stat-cache loss

    Since the fix discussion centers on whether index reuse is worth it, I A/B'd exactly that — identical git add -A -- ., only the index lifecycle differs:

    run 1 run 2
    fresh index + read-tree HEAD (current behavior) 7.96s 6.72s
    one persistent index, reused 6.19s 0.27s

    25x on the steady-state run. The current code pays run-1 cost on every single capture, forever. read-tree HEAD populates the private index with no stat information, so add -A must re-open and re-hash all 21,789 files each turn — and because the index is discarded, no git-side tuning (core.fsmonitor, untrackedCache) can ever engage. This is a direct measurement of what #10792 describes as "starting with a fresh index throws away Git's stat cache on every capture."

    Why it intermittently crosses 30s here

    At ~7s warm I sit ~4x under the limit, so this fails intermittently rather than every turn but it does fail. Contributing factors:

    • ~4s of the 7.3s is non-CPU wait (1.46s user + 1.76s sys), i.e. dominated by file-open latency rather than hashing. On a managed Mac with endpoint security hooking every open, that multiplier is not small.
    • VcsProcess permits 8 concurrent git processes, so multiple threads in the same repo finishing turns together each run their own full-tree re-hash against the same disk.
    • Capture fires at turn end, concurrent with status broadcast and turn-diff computation, right after the agent has churned the page cache.

    tmp_pack litter reproduces on macOS too

    $ git count-objects -v
    warning: garbage found: .git/objects/pack/tmp_pack_ZZjjBK
    count: 547
    in-pack: 43384
    prune-packable: 148
    garbage: 1
    size-garbage: 503
    

    Smaller than the Linux report's 142 × ~800 MiB, but the same mechanism, and nothing cleans it up.

    Note on #11665

    Retrying transient git failures won't help this class of failure and could worsen it. The timeout here is deterministic, not transient that's the "retries a guaranteed 30s timeout" in this issue's title. If #11665 lands before #10792, affected workspaces get more doomed 30s git processes per turn, not fewer. The index-reuse fix needs to come first.

  5. vedprakash2302 commented on Sep 15, 2026

    @vedprakash2302
    Contributor

    Additional Windows desktop / WSL2 reproduction

    What happened

    I use the T3 Code desktop app on Windows with a WSL backend only. The app repeatedly shows this error while working in a large monorepo:

    VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repository-root>) after 30000ms
    

    Diagnosis

    The installed WSL backend still creates a fresh temporary GIT_INDEX_FILE for each capture, runs git read-tree HEAD, then git add -A -- ., followed by write-tree and commit-tree. The Git process runner supplies a 30,000 ms default, and capture does not override it.

    This matches the checkpoint implementation in GitVcsDriver.ts at v0.0.40 and the VcsProcess timeout. The installed nightly bundle was also inspected directly and has the same sequence.

    The fresh index loses the normal index's cached file-stat information, making the capture substantially more expensive than ordinary status. This is the likely bottleneck; the local traces identify the capture operation but do not identify which Git subcommand timed out. I have not independently timed a full temporary-index capture.

    Reproduction conditions

    1. Open a large Git monorepo at its repository root in Windows desktop using the WSL backend.
    2. Use a checkout with 458,319 tracked files, cone-mode sparse checkout, and sparse index enabled.
    3. Run agent turns. The local traces show repeated failures during pre-turn baseline capture and turn-completion capture.

    This is an observed reproduction in an existing private repository, not a minimal public reproducer.

    Version and environment

    • Installed WSL backend bundle: 0.0.41-nightly.20260915.1766.
    • Triage CLI: t3@latest, resolved to 0.0.40 on September 15, 2026. The generated context's version refers to that CLI, not the nightly backend.
    • Windows desktop client, WSL2 Linux x64 backend.
    • Kernel: 6.6.114.1-microsoft-standard-WSL2.
    • Repository filesystem: native WSL ext4, not a Windows-mounted filesystem.
    • Git: 2.52.0.vfs.0.5.
    • Node in triage context: v24.18.0.
    • Agent: OpenCode.

    Evidence

    Representative local trace durations:

    VcsProcess.run:                             30076.546 ms, timeout failure
    GitVcsDriver.checkpoints.captureCheckpoint: 33312.452 ms, failure
    ensurePreTurnBaselineFromDomainTurnStart:   33592.280 ms, failure
    

    Ordinary Git status timings from the same checkout:

    git --no-optional-locks status --porcelain=v1 -uno
    0.173 seconds
    
    git --no-optional-locks status --porcelain=v1 --untracked-files=all
    1.884 seconds
    

    These measurements show that fast ordinary Git status does not rule out this checkpoint failure. The WSL repository already uses native ext4.

    Related issue and workaround

    This appears to be another reproduction of #3646. #10792 proposes reusing Git index metadata and was still open when checked on September 15, 2026. I have not tested that proposed fix.

    No workaround was applied during this triage.

    Triage attribution

    Investigated with OpenCode, model github-copilot/gpt-6-astra, following the t3 triage playbook. The CLI generated its context and prompt; this existing OpenCode session continued the investigation because neither supported standalone triage agent CLI, claude or codex, was on PATH.

  6. mikeohuo commented on Sep 16, 2026

    @mikeohuo
    Author

    Triage follow-up: #10792 is in 0.0.41-nightly.20260916.1795 but never engages on this repo

    What happened

    Turn diffs still report missing on every turn in my Unity monorepo after updating to the nightly that contains #10792. Checkpoint capture still runs the fresh-index git add -A and is killed at the 30 s timeout.

    Diagnosis

    Grounded in the source at tag v0.0.41-nightly.20260916.1795 (commit 3060cc46), which matches the installed bundle.

    captureCheckpoint copies the workspace index, runs read-tree --reset HEAD (apps/server/src/vcs/GitVcsDriver.ts:792), then inspects the result with git ls-files -v capped at WORKSPACE_FILES_MAX_OUTPUT_BYTES = 16 MiB (GitVcsDriver.ts:380, :800). The runner's default outputMode is "error" (apps/server/src/processRunner.ts:291), so a listing over the cap fails with ProcessOutputLimitError. Effect.orElseSucceed(() => false) at GitVcsDriver.ts:806 turns that into "do not reuse", and the driver falls through to the fresh read-tree HEAD at :811 followed by add -A at :820, which is the pre-#10792 behaviour and cannot finish inside DEFAULT_TIMEOUT_MS = 30_000 (VcsProcess.ts:54).

    On this checkout the workspace is the game/ subdirectory of the repo root, ls-files -v returns 268,730 entries and 24.3 MB, and the index has zero assume-unchanged or skip-worktree entries. So the flag check would pass; only the size cap fails.

    Steps to reproduce

    1. Open a workspace in a Git repo where git ls-files -v from the workspace directory exceeds 16 MiB (roughly 250k+ tracked files).
    2. Run any agent turn.
    3. Observe thread.turn-diff-completed events with status: "missing" and a fresh t3-checkpoint-index-<uuid>.lock left in .git about 30 s after each turn boundary.

    Version

    Desktop 0.0.41-nightly.20260916.1795 (Windows installer). Triage CLI t3@nightly, same version.

    Environment

    Windows 11 Pro 10.0.26200, x64. Desktop app against its local server. git 2.45.2.windows.1 with Git LFS (filter.lfs.required=true, 42 LFS patterns). Provider claudeAgent. Claude Code 2.1.273, Codex 0.154.0.

    Evidence

    # Replayed both capture paths against the real workspace with a temporary GIT_INDEX_FILE.
    # Both produce the identical tree 4fcf501f238410a0d7766f3a8e4347cf292479cc.
    fresh index:            read-tree 319 ms | add -A 115361 ms | write-tree 347 ms
    copied index + --reset: read-tree 381 ms | add -A    577 ms | write-tree 173 ms
    
    # The inspection step that gates the fast path:
    git ls-files -v   -> 268730 lines, 24265187 bytes   (cap: 16777216)
    special-flag entries (^[a-zS] ): 0
    
    # Two captures at a new thread's start, both killed at the timeout (lock left behind):
    .git/t3-checkpoint-index-67d5da25-....lock   2026-09-16 07:55:30 UTC   0 bytes
    .git/t3-checkpoint-index-2d0b9255-....lock   2026-09-16 07:56:02 UTC   0 bytes
    
    # state.sqlite, this project: every thread.turn-diff-completed since June is status="missing",
    # except one thread on 2026-09-15 that ran in a small sparse worktree (status="ready").

    Additional litter: the Effect.ensuring(cleanupTempIndex) at GitVcsDriver.ts:864 removes the temp index but not Git's <index>.lock, so each killed capture leaves a 0-byte lock file. This repo has 1,690 of them plus two orphaned ~43 MB temp indexes, on top of the tmp_pack_* files from the original report.

    Related issues

    This issue (#3646) and #10792. Not a duplicate of #10905 (0-byte checkpoint refs) or #9776 (LFS tmp fill from background polling).

    Suggested fix

    Give the ls-files -v inspection its own, much larger cap, or stream it in "truncate" mode and scan only for ^[a-zS] lines, and fall back only when a flagged entry is actually present. With that change this repo would go from a guaranteed timeout to about 1.5 s per capture.

    Fix applied or workaround

    None. No local patch, no database write.

    Filed by

    Claude Code (claude-fable-5-1) via the t3 triage playbook.

  7. mindfield83 commented on Sep 16, 2026

    @mindfield83

    macOS reproduction with a small repository (2,418 tracked files, 77 MB), plus a measurement that points at a cheap fix.

    Environment: T3 Code 0.0.40 desktop, macOS 26.6.2, git 2.50.1 (Apple Git-155). Repository on an external USB-C SSD (APFS inside a sparse bundle), also watched by Syncthing.

    Error: VcsProcessTimeoutError: VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repo>) after 30000ms, logged in server.trace.ndjson on captureCheckpointFromTurnCompletion (durationMs 39236). Two checkpoints three minutes later succeeded in 1.6 s and 1.8 s, so this is not a large-repo problem here; it is an I/O-contention problem.

    Why add -A is expensive at all. captureCheckpoint seeds the temporary index with git read-tree HEAD. An index produced by read-tree carries no stat data, so the following git add -A -- . cannot use the lstat shortcut and has to read and hash every tracked file. On an idle volume that is 0.3 s for this repo; as soon as anything else reads the same volume (a Syncthing full scan, a large copy, a sleeping USB disk) it exceeds 30 s.

    Reproduced outside T3 with the same command sequence (GIT_INDEX_FILE=<tmp> git read-tree HEAD; git add -A -- .; git write-tree), three runs while a large rsync was running on the same SSD:

    Index seeded from read-tree add -A
    read-tree HEAD (current T3 behaviour), idle volume 0.02–0.04 s 0.29–0.32 s
    read-tree HEAD, concurrent I/O, run 1/2/3 13.5 / 3.3 / 5.2 s 135 / 139 / 128 s
    copy of the real index (git rev-parse --git-path index), same concurrent I/O — 1.6 s

    Suggestion (complements #8301): seed the temporary index from a copy of the real index and fall back to read-tree HEAD only when there is none. The stat cache then lets git add -A skip unchanged files with an lstat instead of reading their content; the checkpoint result is identical because add -A still picks up every change relative to the working tree. In the measurement above that is an 80× difference under load, and it keeps the 30 s budget realistic even on slow or contended volumes. A configurable or longer capture timeout (and back-off, as proposed in #4517) would still be worthwhile for the cases in this thread where the working tree itself is huge.

  8. callumenator commented on Sep 16, 2026

    @callumenator

    Adding a data point - in my case the timeouts (and the leftover packs) came from capture running git add -A over large untracked files.

    I had an untracked/unignored db dump and some playwright trace zips in the working tree. Git was writing 512MB (the default core.bigFileThreshold) straight into a pack, so each capture that hit the 30s limit partway through left a tmp_pack_* behind. I got about 1.2GiB every 30 seconds and ended up with roughly 380gb of them (I got down to 2gb free on my Mac before catching it). Killing add -A partway through a large untracked file reproduces it.

    Also, when capture does succeed the checkpoint ref pins the file, and only a revert in t3 (to an earlier checkpoint) clears those refs. My repo has ~2000 checkpoint refs back to April holding 1.8gb, that no branch references.

    #10792 doesn’t (I think) help in this case - an untracked file isn’t in the index to re-use. Cleaning up failed packs (#9809) would stop the disk filling but capture would still time out each turn and still store these files when it succeeds. Skipping large untracked files during capture might be worth considering, also some sort of reaper for old thread checkpoints.

    macOS 26 arm64, git 2.52.0, T3 code 0.0.40

  9. mikeohuo commented on Sep 17, 2026

    @mikeohuo
    Author
    Image

    Thanks everyone and @juliusmarminge!! 🥹 I've escaped the T3Code mono-repo underclass

  10. dimitriosbleh-afk commented on Sep 21, 2026

    @dimitriosbleh-afk

    This still occurred on macOS 26.5.1 with T3 Code Nightly 0.0.43-nightly.20260921.2044 on 21 September 2026. The large-untracked-file case described above appears to remain after the index-reuse fixes.

    In my existing project, one untracked and unignored .tar.gz container-image export was 1,972,946,594 bytes. T3 repeatedly showed:

    VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repository-root>) after 30000ms
    

    The server trace recorded four consecutive capture failures at approximately 30.06–30.11 seconds each. .git/objects/pack/tmp_pack_* files accumulated during the failures. Cleanup ultimately removed 24 abandoned temporary packs totalling 40,433,746,208 bytes, after checking that they were no longer open. The original archive was preserved.

    Workaround and verification: adding only that archive’s path to .git/info/exclude stopped the failures without restarting T3. T3’s own trace then recorded three successful checkpoint captures at 163.277 ms, 175.428 ms and 149.488 ms. A separate replay of the checkpoint Git operations using a temporary index completed in 0.357 seconds, with the real index unchanged. After cleanup, git count-objects -v reported garbage: 0.

    The installed bundle already contains the copied-index / read-tree --reset HEAD path. This observation is from an existing working repository, not a clean synthetic reproduction. Diagnosis and timing checks were performed with Codex.

    Could you reopen this for the remaining large-untracked-file case, or track it separately? An actionable oversized-file message, bounded retries, and cleanup limited to temporary packs owned by a terminated capture would help prevent the repeated disk growth.

  11. chhoumann commented on Oct 3, 2026

    @chhoumann

    Reopen #3646: one large untracked file still makes checkpoint timeouts fill the disk

    Could #3646 be reopened for the large-untracked-file case? On October 3, checkpoint attempts filled my Mac's disk with 66 abandoned tmp_pack_* files occupying 82.83 GB. I then reproduced the failure in a disposable repository with one tracked text file and one untracked binary, using the checkpoint implementation from an official current nightly.

    This is the case described in the September 16 comment and the September 21 follow-up. Reusing the index helps unchanged tracked files, but the large untracked file still has to be read and packed on each unsuccessful capture.

    Version and environment

    • T3 Code 0.0.46-nightly.20261003.2632, commit f391794a35c604d57e166a3ab48d56fc6e4e469a.
    • macOS 27.0, build 26A428, arm64, local APFS storage.
    • Git 2.55.0. The isolated probe uses Node 24.21.0.
    • The disposable repository has one tracked README.md, no remotes, and no Git LFS configuration. Global and system Git configuration are excluded.

    The version above is the verified reproduction version. The original incident spanned a nightly update, so I am not attributing every earlier failure to that exact build.

    Minimal reproduction

    1. Initialize a local Git repository and commit one small README.md.
    2. Create one untracked, unignored 2 GiB file containing data that does not compress well. The fixture repeats a random 1 MiB block, longer than zlib's compression window. A zero-filled or sparse file is not an equivalent test.
    3. Construct the driver with GitVcsDriver.makeVcsDriverShape() and call driver.checkpoints.captureCheckpoint(...) with a new checkpoint ref and the repository root as cwd, using the real VcsProcess, ProcessRunner, and Node services.
    4. After the 30-second timeout, inspect .git/objects/pack/ and run git count-objects -v.
    5. Request another capture without changing the binary. A second abandoned pack remains after another timeout.
    6. Add only /large-untracked.bin to .gitignore and capture again. The capture succeeds, while the two previous temporary packs remain.

    The accompanying reproduction.zip contains a runnable probe and its JSON results:

    python3 reproduce.py --output results.json

    The probe requires the pinned macOS nightly, Python 3.10+, Node 24+, and Git. It needs about 7 GiB of temporary free space. It copies the installed bundle into a temporary directory and appends one export to invoke the existing checkpoint function. The production function bodies, 30-second timeout, and process termination behavior are unchanged; there are no mocked Git commands or injected timeouts. It does not launch a provider, open a T3 profile, or touch an existing repository. Temporary files are removed after the run.

    This is a direct test of the shipped checkpoint component, not a separate replay through the desktop UI and provider orchestration. The original incident below supplies evidence from normal application use. On faster hardware, increase the fixture with --size-mib 4096 if 2 GiB finishes within the deadline.

    Observed results

    Capture Result Duration Abandoned packs after return
    Small repository control Succeeded 77.2 ms 0
    With the untracked 2 GiB file Timed out 30,041.4 ms 1
    Same file, second capture Timed out 30,047.4 ms 2
    Same file now ignored Succeeded 81.0 ms 2

    The first two failed captures left 3,457,425,408 allocated bytes of temporary packs in total. The failure is:

    VcsProcessTimeoutError: VCS process timed out in
    GitVcsDriver.checkpoints.captureCheckpoint: git (<repo>) after 30000ms
    

    Only the successful control captures created checkpoint refs. The real Git index and HEAD remained unchanged, and the source binary remained intact. git count-objects -v classified the leftover packs as garbage. A successful later capture did not remove them.

    Impact observed in normal use

    A 1,563,981,564-byte video was downloaded into my existing repository at 13:28 CEST. Temporary packs began appearing at 13:32. By approximately 16:20, 66 temporary packs occupied 82,829,254,656 allocated bytes. The disk had roughly 120 MB free, and ordinary file writes were failing with ENOSPC.

    T3's server trace recorded repeated checkpoint timeouts, including these separate capture attempts:

    2026-10-03T14:12:36.681831Z  durationMs=30056.089083
    2026-10-03T14:13:06.789825Z  durationMs=30049.452792
    2026-10-03T14:25:47.895864Z  durationMs=30055.357625
    
    VcsProcessTimeoutError: VCS process timed out in
    GitVcsDriver.checkpoints.captureCheckpoint: git (<repo>) after 30000ms
    

    I also observed the T3 server spawning:

    git -C <repo> \
      -c core.fsmonitor=false \
      -c sparse.expectFilesOutsideOfPatterns=false \
      -c core.fsync=objects,reference \
      -c core.fsyncMethod=fsync \
      add -A -- .

    Excluding the downloaded source media and removing verified-unused temporary files recovered the space. T3 then created a real checkpoint containing the analysis note and excluding those media files. git fsck --connectivity-only --no-dangling passed, and git count-objects -v reported zero garbage. Source media was preserved; cleanup targeted abandoned temporary files only.

    Why this survives the index-reuse fix

    At the verified revision:

    1. captureCheckpoint isolates the index, but does not isolate new object writes. Git still writes to the shared repository object store.
    2. Staging runs git add -A -- . without a timeout override. VcsProcess supplies 30,000 ms.
    3. The finalizer removes the temporary index and its lock, not the temporary pack left by the terminated Git process.
    4. A later capture sees the same untracked binary and repeats the expensive write. These are successive capture requests; the evidence does not depend on claiming that VcsProcess internally retries timeout errors.

    This reproduction involves no fetch, so it is separate from the background-fetch path in #3525. It also does not require accumulated worktrees or retention of completed checkpoint objects.

    Expected behavior and regression coverage

    A failed checkpoint must not make repeated use of the application consume unbounded disk space. Increasing the timeout alone would not address leftover storage after cancellation or another timeout.

    A focused fix should be verifiable with these cases:

    • A timeout or cancellation after Git begins a pack leaves no storage owned by that failed capture after the child process has exited.
    • Repeating a capture over the unchanged failing input does not accumulate temporary packs indefinitely.
    • Cleanup preserves valid objects, existing checkpoint refs, the user's real index, and temporary files belonging to concurrent Git operations. A blanket deletion of shared tmp_pack_* files would not be safe.
    • Successful captures still preserve the intended file contents. If an oversized-file policy excludes files, the user can see that they are outside checkpoint/restore coverage.

    Capture-owned temporary object storage is one possible approach. #9809 proposed related cleanup, but it is closed and unmerged; I have not tested that patch and am not assuming it is the required implementation.

    The immediate workaround is to exclude the large source file from Git and then remove abandoned temporary files only after confirming they are unused. The file remains on disk but is no longer covered by checkpoints.

    Investigation and reproduction were performed with Codex. The attached results distinguish the live incident from the isolated component test.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions