Repository navigation
[Bug]: Checkpoint capture on large monorepos retries a guaranteed 30s git-add timeout every turn — permanent CPU burn + tmp_pack disk litter #3646
Description
Activity
Additional Linux desktop reproduction:
- I confirmed that the T3 Code server process directly spawns git -C add -A -- .; the application, not the target project, owns the Git child.
- The workspace had one unignored ELF core dump, 32 GiB apparent size and 4.3 GiB allocated. That is enough to make every automatic checkpoint attempt expensive.
- While observing the failure, Git garbage grew from 132 to 142 tmp_pack_* files. The new files were about 790-803 MiB each, with final mtimes spaced almost exactly 30 seconds apart.
- Current source still has the same chain: CheckpointReactor captures a pre-turn baseline from message-sent, turn-start-requested, and turn.started events; GitVcsDriver runs add -A -- .; VcsProcess supplies the 30,000 ms default timeout; and capture cleanup removes only the temporary index. I found no source-side cleanup for stale tmp_pack_* files.
I did not capture the server timeout log itself, but the live process ancestry, exact command, source path, and 30-second cadence make the timeout/retry explanation very strong. This reproduces outside a large monorepo too: a single large, unignored crash dump is sufficient.
An adaptive timeout alone would still put crash dumps into hidden checkpoint history. A safer resolution would make capture failure back off for an unchanged workspace, expose an actionable reason, and clean only temp packs created by a fully terminated failed capture.
Reacted by Michael Huo and Mohamed HosnieFollow-up reproduction on an official nightly
This still reproduces when the Git root, T3 project directory, and backend working directory are all the same directory. The agent completes its work, but its conversation turn has no usable checkpoint diff.
Environment and scope
- Windows; a large repository on a secondary drive.
- Official installed nightly:
0.0.39-nightly.20260906.1316. - The unmodified installed server bundle ran with an isolated T3 profile and its working directory set to the repository root.
- The T3 project used the same repository root, with
.gitdirectly inside it. - The checkout already contained unrelated uncommitted changes and a linked worktree directory. This was not a clean, minimal reproduction.
Observed result
A
claudeAgentturn completed successfully and created the requested small text file. Existing unrelated worktree changes were also present.Native
vcs.refreshStatus,review.getDiffPreview, and file-content requests returned data. The proof file's ordinary diff contained the exact expected text. These were backend/provider checks; I did not visually verify the rendered UI.For this one turn, the server logged three separate checkpoint timeouts. Each reported the following error, with the repository path redacted:
VcsProcessTimeoutError: VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repository-root>) after 30000msThe next warning was:
checkpoint capture missing pre-turn baselineNo checkpoint ref was created for the test thread. Its checkpoint list remained empty.
orchestration.getTurnDifffailed withCheckpointTurnRangeUnavailableError, including:Turn diff range exceeds current turn count: requested 1, current 0.Four empty temporary checkpoint-index
.lockfiles remained afterward. This run did not measure CPU utilization or observe newtmp_pack_*files; those remain separate observations from the earlier reports.What this separates
The checkpoint timeout persists with all three directories aligned at the Git root. The separate review-directory validation issue was not blocking these requests.
The practical failure here is missing per-turn conversation diffs after successful agent edits. Ordinary Git diff access does not establish that checkpoint-based diffs work.
Proposed next steps
Please prioritize successful checkpoint capture on large dirty repositories, with bounded retries and cleanup as supporting fixes.
Several existing PRs address different parts of this:
- PR 8301: stage only changed checkpoint paths targets the expensive capture operation. It looks directly relevant, but currently conflicts with main and needs review and rebasing. I have not tested it.
- PR 4517: back off after repeated capture failures reduces repeated unsuccessful work. It does not itself make missing turn diffs available.
- PR 9809: clean up failed checkpoint writes addresses leftover temporary storage. It retains the 30-second deadline; its description also notes that native Windows was not tested.
Useful acceptance coverage:
- Test a large dirty repository with both root and nested project directories.
- Cover staged, unstaged, deleted, and untracked files, including a newly staged file subsequently deleted from disk.
- Confirm that two consecutive turns create checkpoints and return the exact changes for each turn.
- Confirm that checkpoint operations preserve the user's real Git index and unrelated files.
- On failure, bound repeated attempts, clean only the failed capture's own temporary storage, and show a useful failure reason.
I can retest an official nightly containing a fix against this repository.
Audit attribution
This audit was performed by Astra and Claude Fable 5.1 agents. Astra performed the source inspection and native endpoint checks. Fable 5.1 executed the live agent smoke test and independently reviewed the findings.
Additional Linux/NFS reproduction on T3 Code 0.0.40:
Environment
- T3 server running on a remote Linux machine
- client connected remotely
- repository stored on an NFSv4 volume
- approximately 28,000 tracked files
.gitdirectory approximately 2.6 GB
Observed behavior
An otherwise successful agent turn ended with:
VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (
<workspace>) after 30000msThe response and working-tree changes remained available. The failure therefore appeared non-blocking, but the checkpoint for that turn was unavailable and the associated diff/revert may be incomplete.
There was no Git lock, and the repository remained usable.
git status --porcelaincompleted in approximately 2.3 seconds.Timings
I reproduced the checkpoint sequence with a temporary
GIT_INDEX_FILEand timed its main operations independently:git read-tree HEAD: approximately 0.8 secondsgit add -A -- .: approximately 21–24 seconds without heavy contentiongit write-tree: approximately 0.3 seconds
At the time of the original failure, several I/O-heavy transfers were using the same shared NFS volume. The
git add -Astep was already close enough to the fixed 30-second deadline that this additional contention appears sufficient to explain the timeout.Temporary checkpoint index files were cleaned up correctly in this reproduction, and I found no indication of repository corruption.
Limited comparison
While investigating, I briefly inspected the public
openai/codexrepository. Its per-turn diff tracker appears to retain exact deltas from recordedapply_patchmutations without rereading the entire workspace. This is only a limited observation, not a suggested implementation: I did not study all of its editing and restoration paths, its desktop client is not fully represented in the public repository, and arbitrary shell edits involve different tradeoffs.I am mentioning it only as evidence that change-tracking designs can have different performance characteristics. The appropriate checkpoint semantics for T3 are best left to the maintainers.
Additional macOS reproduction on Apple Silicon. Notable because the repo is ordinary.
The existing reports all involve something extreme: a Unity monorepo, a 32 GiB unignored core dump, an NFSv4 volume. This one is a plain local-APFS checkout of
pingdotgg/t3codeitself, and it still fails.Environment
- macOS, Darwin 25.6.0, arm64, local APFS SSD
git version 2.50.1 (Apple Git-155)- Repository:
pingdotgg/t3code@main, 21,789 tracked files,.git124 MB - Workspace cwd is the repo root; no linked worktrees, no hooks, no
core.fsmonitor/untrackedCache/preloadindexset git status --porcelainreturns in well under a second; no staleindex.lock
Timings
Reproduced the capture sequence against a temporary
GIT_INDEX_FILE, warm cache, machine otherwise idle:step real git read-tree HEAD0.11s git add -A -- .6.7–8.0s git write-tree0.08s Same shape as the NFS report above (0.8s / 21–24s / 0.3s), just proportionally smaller.
Isolating the stat-cache loss
Since the fix discussion centers on whether index reuse is worth it, I A/B'd exactly that — identical
git add -A -- ., only the index lifecycle differs:run 1 run 2 fresh index + read-tree HEAD(current behavior)7.96s 6.72s one persistent index, reused 6.19s 0.27s 25x on the steady-state run. The current code pays run-1 cost on every single capture, forever.
read-tree HEADpopulates the private index with no stat information, soadd -Amust re-open and re-hash all 21,789 files each turn — and because the index is discarded, no git-side tuning (core.fsmonitor,untrackedCache) can ever engage. This is a direct measurement of what #10792 describes as "starting with a fresh index throws away Git's stat cache on every capture."Why it intermittently crosses 30s here
At ~7s warm I sit ~4x under the limit, so this fails intermittently rather than every turn but it does fail. Contributing factors:
- ~4s of the 7.3s is non-CPU wait (1.46s user + 1.76s sys), i.e. dominated by file-open latency rather than hashing. On a managed Mac with endpoint security hooking every open, that multiplier is not small.
VcsProcesspermits 8 concurrent git processes, so multiple threads in the same repo finishing turns together each run their own full-tree re-hash against the same disk.- Capture fires at turn end, concurrent with status broadcast and turn-diff computation, right after the agent has churned the page cache.
tmp_pack litter reproduces on macOS too
$ git count-objects -v warning: garbage found: .git/objects/pack/tmp_pack_ZZjjBK count: 547 in-pack: 43384 prune-packable: 148 garbage: 1 size-garbage: 503Smaller than the Linux report's 142 × ~800 MiB, but the same mechanism, and nothing cleans it up.
Note on #11665
Retrying transient git failures won't help this class of failure and could worsen it. The timeout here is deterministic, not transient that's the "retries a guaranteed 30s timeout" in this issue's title. If #11665 lands before #10792, affected workspaces get more doomed 30s git processes per turn, not fewer. The index-reuse fix needs to come first.
Additional Windows desktop / WSL2 reproduction
What happened
I use the T3 Code desktop app on Windows with a WSL backend only. The app repeatedly shows this error while working in a large monorepo:
VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repository-root>) after 30000msDiagnosis
The installed WSL backend still creates a fresh temporary
GIT_INDEX_FILEfor each capture, runsgit read-tree HEAD, thengit add -A -- ., followed bywrite-treeandcommit-tree. The Git process runner supplies a 30,000 ms default, and capture does not override it.This matches the checkpoint implementation in GitVcsDriver.ts at v0.0.40 and the VcsProcess timeout. The installed nightly bundle was also inspected directly and has the same sequence.
The fresh index loses the normal index's cached file-stat information, making the capture substantially more expensive than ordinary status. This is the likely bottleneck; the local traces identify the capture operation but do not identify which Git subcommand timed out. I have not independently timed a full temporary-index capture.
Reproduction conditions
- Open a large Git monorepo at its repository root in Windows desktop using the WSL backend.
- Use a checkout with 458,319 tracked files, cone-mode sparse checkout, and sparse index enabled.
- Run agent turns. The local traces show repeated failures during pre-turn baseline capture and turn-completion capture.
This is an observed reproduction in an existing private repository, not a minimal public reproducer.
Version and environment
- Installed WSL backend bundle:
0.0.41-nightly.20260915.1766. - Triage CLI:
t3@latest, resolved to0.0.40on September 15, 2026. The generated context's version refers to that CLI, not the nightly backend. - Windows desktop client, WSL2 Linux x64 backend.
- Kernel:
6.6.114.1-microsoft-standard-WSL2. - Repository filesystem: native WSL ext4, not a Windows-mounted filesystem.
- Git:
2.52.0.vfs.0.5. - Node in triage context:
v24.18.0. - Agent: OpenCode.
Evidence
Representative local trace durations:
VcsProcess.run: 30076.546 ms, timeout failure GitVcsDriver.checkpoints.captureCheckpoint: 33312.452 ms, failure ensurePreTurnBaselineFromDomainTurnStart: 33592.280 ms, failureOrdinary Git status timings from the same checkout:
git --no-optional-locks status --porcelain=v1 -uno 0.173 seconds git --no-optional-locks status --porcelain=v1 --untracked-files=all 1.884 secondsThese measurements show that fast ordinary Git status does not rule out this checkpoint failure. The WSL repository already uses native ext4.
Related issue and workaround
This appears to be another reproduction of #3646. #10792 proposes reusing Git index metadata and was still open when checked on September 15, 2026. I have not tested that proposed fix.
No workaround was applied during this triage.
Triage attribution
Investigated with OpenCode, model
github-copilot/gpt-6-astra, following thet3 triageplaybook. The CLI generated its context and prompt; this existing OpenCode session continued the investigation because neither supported standalone triage agent CLI,claudeorcodex, was on PATH.Triage follow-up: #10792 is in 0.0.41-nightly.20260916.1795 but never engages on this repo
What happened
Turn diffs still report
missingon every turn in my Unity monorepo after updating to the nightly that contains #10792. Checkpoint capture still runs the fresh-indexgit add -Aand is killed at the 30 s timeout.Diagnosis
Grounded in the source at tag
v0.0.41-nightly.20260916.1795(commit3060cc46), which matches the installed bundle.captureCheckpointcopies the workspace index, runsread-tree --reset HEAD(apps/server/src/vcs/GitVcsDriver.ts:792), then inspects the result withgit ls-files -vcapped atWORKSPACE_FILES_MAX_OUTPUT_BYTES= 16 MiB (GitVcsDriver.ts:380,:800). The runner's defaultoutputModeis"error"(apps/server/src/processRunner.ts:291), so a listing over the cap fails withProcessOutputLimitError.Effect.orElseSucceed(() => false)atGitVcsDriver.ts:806turns that into "do not reuse", and the driver falls through to the freshread-tree HEADat:811followed byadd -Aat:820, which is the pre-#10792 behaviour and cannot finish insideDEFAULT_TIMEOUT_MS = 30_000(VcsProcess.ts:54).On this checkout the workspace is the
game/subdirectory of the repo root,ls-files -vreturns 268,730 entries and 24.3 MB, and the index has zero assume-unchanged or skip-worktree entries. So the flag check would pass; only the size cap fails.Steps to reproduce
- Open a workspace in a Git repo where
git ls-files -vfrom the workspace directory exceeds 16 MiB (roughly 250k+ tracked files). - Run any agent turn.
- Observe
thread.turn-diff-completedevents withstatus: "missing"and a fresht3-checkpoint-index-<uuid>.lockleft in.gitabout 30 s after each turn boundary.
Version
Desktop
0.0.41-nightly.20260916.1795(Windows installer). Triage CLIt3@nightly, same version.Environment
Windows 11 Pro 10.0.26200, x64. Desktop app against its local server. git 2.45.2.windows.1 with Git LFS (
filter.lfs.required=true, 42 LFS patterns). ProviderclaudeAgent. Claude Code 2.1.273, Codex 0.154.0.Evidence
# Replayed both capture paths against the real workspace with a temporary GIT_INDEX_FILE. # Both produce the identical tree 4fcf501f238410a0d7766f3a8e4347cf292479cc. fresh index: read-tree 319 ms | add -A 115361 ms | write-tree 347 ms copied index + --reset: read-tree 381 ms | add -A 577 ms | write-tree 173 ms # The inspection step that gates the fast path: git ls-files -v -> 268730 lines, 24265187 bytes (cap: 16777216) special-flag entries (^[a-zS] ): 0 # Two captures at a new thread's start, both killed at the timeout (lock left behind): .git/t3-checkpoint-index-67d5da25-....lock 2026-09-16 07:55:30 UTC 0 bytes .git/t3-checkpoint-index-2d0b9255-....lock 2026-09-16 07:56:02 UTC 0 bytes # state.sqlite, this project: every thread.turn-diff-completed since June is status="missing", # except one thread on 2026-09-15 that ran in a small sparse worktree (status="ready").
Additional litter: the
Effect.ensuring(cleanupTempIndex)atGitVcsDriver.ts:864removes the temp index but not Git's<index>.lock, so each killed capture leaves a 0-byte lock file. This repo has 1,690 of them plus two orphaned ~43 MB temp indexes, on top of thetmp_pack_*files from the original report.Related issues
This issue (#3646) and #10792. Not a duplicate of #10905 (0-byte checkpoint refs) or #9776 (LFS tmp fill from background polling).
Suggested fix
Give the
ls-files -vinspection its own, much larger cap, or stream it in"truncate"mode and scan only for^[a-zS]lines, and fall back only when a flagged entry is actually present. With that change this repo would go from a guaranteed timeout to about 1.5 s per capture.Fix applied or workaround
None. No local patch, no database write.
Filed by
Claude Code (claude-fable-5-1) via the
t3 triageplaybook.- Open a workspace in a Git repo where
macOS reproduction with a small repository (2,418 tracked files, 77 MB), plus a measurement that points at a cheap fix.
Environment: T3 Code 0.0.40 desktop, macOS 26.6.2, git 2.50.1 (Apple Git-155). Repository on an external USB-C SSD (APFS inside a sparse bundle), also watched by Syncthing.
Error:
VcsProcessTimeoutError: VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repo>) after 30000ms, logged inserver.trace.ndjsononcaptureCheckpointFromTurnCompletion(durationMs 39236). Two checkpoints three minutes later succeeded in 1.6 s and 1.8 s, so this is not a large-repo problem here; it is an I/O-contention problem.Why
add -Ais expensive at all.captureCheckpointseeds the temporary index withgit read-tree HEAD. An index produced byread-treecarries no stat data, so the followinggit add -A -- .cannot use the lstat shortcut and has to read and hash every tracked file. On an idle volume that is 0.3 s for this repo; as soon as anything else reads the same volume (a Syncthing full scan, a large copy, a sleeping USB disk) it exceeds 30 s.Reproduced outside T3 with the same command sequence (
GIT_INDEX_FILE=<tmp> git read-tree HEAD; git add -A -- .; git write-tree), three runs while a largersyncwas running on the same SSD:Index seeded from read-tree add -A read-tree HEAD(current T3 behaviour), idle volume0.02–0.04 s 0.29–0.32 s read-tree HEAD, concurrent I/O, run 1/2/313.5 / 3.3 / 5.2 s 135 / 139 / 128 s copy of the real index ( git rev-parse --git-path index), same concurrent I/O— 1.6 s Suggestion (complements #8301): seed the temporary index from a copy of the real index and fall back to
read-tree HEADonly when there is none. The stat cache then letsgit add -Askip unchanged files with an lstat instead of reading their content; the checkpoint result is identical becauseadd -Astill picks up every change relative to the working tree. In the measurement above that is an 80× difference under load, and it keeps the 30 s budget realistic even on slow or contended volumes. A configurable or longer capture timeout (and back-off, as proposed in #4517) would still be worthwhile for the cases in this thread where the working tree itself is huge.Adding a data point - in my case the timeouts (and the leftover packs) came from capture running
git add -Aover large untracked files.I had an untracked/unignored db dump and some playwright trace zips in the working tree. Git was writing 512MB (the default core.bigFileThreshold) straight into a pack, so each capture that hit the 30s limit partway through left a tmp_pack_* behind. I got about 1.2GiB every 30 seconds and ended up with roughly 380gb of them (I got down to 2gb free on my Mac before catching it). Killing
add -Apartway through a large untracked file reproduces it.Also, when capture does succeed the checkpoint ref pins the file, and only a revert in t3 (to an earlier checkpoint) clears those refs. My repo has ~2000 checkpoint refs back to April holding 1.8gb, that no branch references.
#10792 doesn’t (I think) help in this case - an untracked file isn’t in the index to re-use. Cleaning up failed packs (#9809) would stop the disk filling but capture would still time out each turn and still store these files when it succeeds. Skipping large untracked files during capture might be worth considering, also some sort of reaper for old thread checkpoints.
macOS 26 arm64, git 2.52.0, T3 code 0.0.40
Reacted by Michael Huo
Thanks everyone and @juliusmarminge!! 🥹 I've escaped the T3Code mono-repo underclass
This still occurred on macOS 26.5.1 with T3 Code Nightly
0.0.43-nightly.20260921.2044on 21 September 2026. The large-untracked-file case described above appears to remain after the index-reuse fixes.In my existing project, one untracked and unignored
.tar.gzcontainer-image export was 1,972,946,594 bytes. T3 repeatedly showed:VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repository-root>) after 30000msThe server trace recorded four consecutive capture failures at approximately 30.06–30.11 seconds each.
.git/objects/pack/tmp_pack_*files accumulated during the failures. Cleanup ultimately removed 24 abandoned temporary packs totalling 40,433,746,208 bytes, after checking that they were no longer open. The original archive was preserved.Workaround and verification: adding only that archive’s path to
.git/info/excludestopped the failures without restarting T3. T3’s own trace then recorded three successful checkpoint captures at 163.277 ms, 175.428 ms and 149.488 ms. A separate replay of the checkpoint Git operations using a temporary index completed in 0.357 seconds, with the real index unchanged. After cleanup,git count-objects -vreportedgarbage: 0.The installed bundle already contains the copied-index /
read-tree --reset HEADpath. This observation is from an existing working repository, not a clean synthetic reproduction. Diagnosis and timing checks were performed with Codex.Could you reopen this for the remaining large-untracked-file case, or track it separately? An actionable oversized-file message, bounded retries, and cleanup limited to temporary packs owned by a terminated capture would help prevent the repeated disk growth.
Reopen #3646: one large untracked file still makes checkpoint timeouts fill the disk
Could #3646 be reopened for the large-untracked-file case? On October 3, checkpoint attempts filled my Mac's disk with 66 abandoned
tmp_pack_*files occupying 82.83 GB. I then reproduced the failure in a disposable repository with one tracked text file and one untracked binary, using the checkpoint implementation from an official current nightly.This is the case described in the September 16 comment and the September 21 follow-up. Reusing the index helps unchanged tracked files, but the large untracked file still has to be read and packed on each unsuccessful capture.
Version and environment
- T3 Code
0.0.46-nightly.20261003.2632, commitf391794a35c604d57e166a3ab48d56fc6e4e469a. - macOS 27.0, build
26A428, arm64, local APFS storage. - Git 2.55.0. The isolated probe uses Node 24.21.0.
- The disposable repository has one tracked
README.md, no remotes, and no Git LFS configuration. Global and system Git configuration are excluded.
The version above is the verified reproduction version. The original incident spanned a nightly update, so I am not attributing every earlier failure to that exact build.
Minimal reproduction
- Initialize a local Git repository and commit one small
README.md. - Create one untracked, unignored 2 GiB file containing data that does not compress well. The fixture repeats a random 1 MiB block, longer than zlib's compression window. A zero-filled or sparse file is not an equivalent test.
- Construct the driver with
GitVcsDriver.makeVcsDriverShape()and calldriver.checkpoints.captureCheckpoint(...)with a new checkpoint ref and the repository root ascwd, using the realVcsProcess,ProcessRunner, and Node services. - After the 30-second timeout, inspect
.git/objects/pack/and rungit count-objects -v. - Request another capture without changing the binary. A second abandoned pack remains after another timeout.
- Add only
/large-untracked.binto.gitignoreand capture again. The capture succeeds, while the two previous temporary packs remain.
The accompanying reproduction.zip contains a runnable probe and its JSON results:
python3 reproduce.py --output results.json
The probe requires the pinned macOS nightly, Python 3.10+, Node 24+, and Git. It needs about 7 GiB of temporary free space. It copies the installed bundle into a temporary directory and appends one export to invoke the existing checkpoint function. The production function bodies, 30-second timeout, and process termination behavior are unchanged; there are no mocked Git commands or injected timeouts. It does not launch a provider, open a T3 profile, or touch an existing repository. Temporary files are removed after the run.
This is a direct test of the shipped checkpoint component, not a separate replay through the desktop UI and provider orchestration. The original incident below supplies evidence from normal application use. On faster hardware, increase the fixture with
--size-mib 4096if 2 GiB finishes within the deadline.Observed results
Capture Result Duration Abandoned packs after return Small repository control Succeeded 77.2 ms 0 With the untracked 2 GiB file Timed out 30,041.4 ms 1 Same file, second capture Timed out 30,047.4 ms 2 Same file now ignored Succeeded 81.0 ms 2 The first two failed captures left 3,457,425,408 allocated bytes of temporary packs in total. The failure is:
VcsProcessTimeoutError: VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repo>) after 30000msOnly the successful control captures created checkpoint refs. The real Git index and
HEADremained unchanged, and the source binary remained intact.git count-objects -vclassified the leftover packs as garbage. A successful later capture did not remove them.Impact observed in normal use
A 1,563,981,564-byte video was downloaded into my existing repository at 13:28 CEST. Temporary packs began appearing at 13:32. By approximately 16:20, 66 temporary packs occupied 82,829,254,656 allocated bytes. The disk had roughly 120 MB free, and ordinary file writes were failing with
ENOSPC.T3's server trace recorded repeated checkpoint timeouts, including these separate capture attempts:
2026-10-03T14:12:36.681831Z durationMs=30056.089083 2026-10-03T14:13:06.789825Z durationMs=30049.452792 2026-10-03T14:25:47.895864Z durationMs=30055.357625 VcsProcessTimeoutError: VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint: git (<repo>) after 30000msI also observed the T3 server spawning:
git -C <repo> \ -c core.fsmonitor=false \ -c sparse.expectFilesOutsideOfPatterns=false \ -c core.fsync=objects,reference \ -c core.fsyncMethod=fsync \ add -A -- .
Excluding the downloaded source media and removing verified-unused temporary files recovered the space. T3 then created a real checkpoint containing the analysis note and excluding those media files.
git fsck --connectivity-only --no-danglingpassed, andgit count-objects -vreported zero garbage. Source media was preserved; cleanup targeted abandoned temporary files only.Why this survives the index-reuse fix
At the verified revision:
captureCheckpointisolates the index, but does not isolate new object writes. Git still writes to the shared repository object store.- Staging runs
git add -A -- .without a timeout override.VcsProcesssupplies 30,000 ms. - The finalizer removes the temporary index and its lock, not the temporary pack left by the terminated Git process.
- A later capture sees the same untracked binary and repeats the expensive write. These are successive capture requests; the evidence does not depend on claiming that
VcsProcessinternally retries timeout errors.
This reproduction involves no fetch, so it is separate from the background-fetch path in #3525. It also does not require accumulated worktrees or retention of completed checkpoint objects.
Expected behavior and regression coverage
A failed checkpoint must not make repeated use of the application consume unbounded disk space. Increasing the timeout alone would not address leftover storage after cancellation or another timeout.
A focused fix should be verifiable with these cases:
- A timeout or cancellation after Git begins a pack leaves no storage owned by that failed capture after the child process has exited.
- Repeating a capture over the unchanged failing input does not accumulate temporary packs indefinitely.
- Cleanup preserves valid objects, existing checkpoint refs, the user's real index, and temporary files belonging to concurrent Git operations. A blanket deletion of shared
tmp_pack_*files would not be safe. - Successful captures still preserve the intended file contents. If an oversized-file policy excludes files, the user can see that they are outside checkpoint/restore coverage.
Capture-owned temporary object storage is one possible approach. #9809 proposed related cleanup, but it is closed and unmerged; I have not tested that patch and am not assuming it is the required implementation.
The immediate workaround is to exclude the large source file from Git and then remove abandoned temporary files only after confirming they are unused. The file remains on disk but is no longer covered by checkpoints.
Investigation and reproduction were performed with Codex. The attached results distinguish the live incident from the isolated component test.
- T3 Code
Before submitting
Area
apps/server
Steps to reproduce
CheckpointReactor→GitVcsDriver.checkpoints.captureCheckpoint, which runsgit read-tree HEAD+git add -A -- .against a tempGIT_INDEX_FILE..git/objects/pack/.Expected behavior
Checkpoint capture should either succeed, or fail once and back off / disable itself for the workspace. A capture that cannot ever succeed should not be retried indefinitely, and killed git processes should not leave permanent garbage in
.git.Actual behavior
VcsProcesshasDEFAULT_TIMEOUT_MS = 30_000andcaptureCheckpointpasses notimeoutMsoverride. On a large monorepo a full-treegit add -Atakes well over 30 seconds, so the process is killed at the timeout every single time. The checkpoint is recorded asmissing(everythread.turn-diff-completedevent in mystate.sqlitehasstatus="missing", across every thread; norefs/t3/checkpoints/*ref has ever been created), and the next turn simply retries. Net effect:core.bigFileThreshold(files that size are common in Unity repos) leaves an orphanedtmp_pack_*in.git/objects/pack/. These accumulate over time —git count-objects -vflags them as garbage, along with a large number of unreferenced loose objects from the partial adds.Related issues — this seems to be the connecting root cause:
Suggested direction (any one of these would resolve the pathology): a per-project or global setting to disable checkpointing; skip-with-backoff after N consecutive capture failures instead of retrying forever; an adaptive/longer timeout; and cleanup of
tmp_pack_*left by killed VCS processes.Impact
Major degradation or frequent failure
Version or commit
0.0.29-nightly.20260701.697 (desktop, Windows installer)
Environment
Windows 11 Pro (10.0.26200), T3 Code Nightly desktop, git 2.54.0.windows.1, provider claudeAgent
Logs or stack traces
Workaround
None found — there's no setting that disables checkpointing. Locally I'm considering patching
captureCheckpointin the bundledbin.mjsto early-exit for this repo, but a nightly update overwrites it.