Skip to content

Checkpoint index reuse falls back on large sparse checkouts and still fails after #10792 #12080

Description

@vedprakash2302

What happened

T3 Code's Windows desktop app with a WSL2 backend repeatedly fails to capture checkpoints in a large sparse-checkout monorepo. The failure persists after updating to a nightly containing #10792:

VCS process timed out in GitVcsDriver.checkpoints.captureCheckpoint:
git (<repository-root>) after 30000ms

Diagnosis

I reproduced the installed backend's checkpoint preparation using a temporary index and isolated Git object storage.

After copying the real index and running:

git -c core.fsmonitor=false read-tree --reset HEAD

the temporary index contained:

  • 458,319 tracked entries.
  • 211,769 skip-worktree entries.
  • 40,621,927 bytes of git ls-files -v output.

The merged optimization accepts index reuse only when the output is not truncated and contains no special flags:

!entries.stdoutTruncated && !/^[a-zS] /m.test(entries.stdout)

The inspection limit is 16,777,216 bytes. Both the listing size and the surviving skip-worktree flags reject reuse in this reproduction.

The fallback's git read-tree HEAD then left all 458,319 entries marked H. The following git add -A -- . took 31.1 seconds, exceeding T3's unchanged 30-second deadline.

With a longer diagnostic timeout, that command finished with exit code 1 because six untracked files existed outside the sparse-checkout definition. A longer timeout alone would therefore not make checkpoint capture succeed in this checkout's current state.

Steps to reproduce

  1. Open a large cone-mode sparse checkout with sparse index enabled.
  2. Have untracked files present outside the sparse-checkout definition.
  3. Run agent turns using the Windows desktop app and WSL backend.
  4. Observe repeated checkpoint timeout errors.

The diagnostic reproduction followed the installed backend's index-copy, reset, inspection, fallback, and add sequence. This is an existing private-repository reproduction, not a minimal public fixture.

Version

0.0.41-nightly.20260916.1795, containing #10792.

Environment

  • Windows desktop client with WSL2 Linux x64 backend.
  • Kernel 6.6.114.1-microsoft-standard-WSL2.
  • Repository on native WSL ext4.
  • Git 2.52.0.vfs.0.5.
  • OpenCode agent.

Evidence

read-tree --reset HEAD: 1.666 seconds, exit 0
ls-files -v:           2.618 seconds, exit 0
index reuse condition: false
fallback read-tree:    2.070 seconds, exit 0
add -A -- .:          31.089 seconds, exit 1

Redacted Git error:

The following paths and/or pathspecs matched paths that exist
outside of your sparse-checkout definition, so will not be
updated in the index:
<six untracked artifact files>

Ordinary status previously completed in 0.173 seconds for tracked changes and 1.884 seconds including untracked files.

The real index remained byte-for-byte unchanged throughout the diagnostic probe.

Related issues

Related to #3646 and its fix #10792. This report isolates the confirmed index-reuse fallback on a large sparse checkout and the sparse-path error revealed when git add is allowed to finish.

Fix applied or workaround

None. The investigation used temporary index and object storage without changing working files or the real staging area.

Filed by

OpenCode, model github-copilot/gpt-6-astra, following the t3 triage playbook.

Activity

  1. juliusmarminge commented on Sep 16, 2026

    @juliusmarminge
    Member

    Triage: real server checkpoint bug. Confirmed on current main and on your 0.0.41-nightly.20260916.1795 (both include #10792). Shipping 0.0.42 does not include #10792 and is worse (always takes the fresh-index path). Not fixed by #10792, and not a wash of #3646.

    What you hit. On a large cone-mode sparse checkout, capture copies the index, runs read-tree --reset HEAD, then inspects git ls-files -v. Reuse only runs when that listing is under 16 MiB and has no assume-unchanged / skip-worktree flags (!entries.stdoutTruncated && !/^[a-zS] /m.test(...)). Your probe had ~40 MiB of listing and ~212k S entries — reuse rejected twice. Fallback read-tree HEAD left every entry H, then git add -A -- . hit the unchanged 30s deadline (31.1s). With a longer timeout the same add exits 1 because untracked files exist outside the sparse cone.

    Why #10792 does not save this. That PR intentionally falls back when S / lowercase flags remain or the listing is truncated — it was protecting manual update-index --skip-worktree / assume-unchanged. Cone-mode sparse-checkout is those S bits. So sparse monorepos never reuse, and the fallback expands the full tree.

    The add is also wrong, not just slow. git add -A -- . exits 1 on outside-cone untracked paths (reproduced on current Git for both workspace and fallback temp indexes). Raising the timeout alone would not make capture succeed while those artifacts exist. Fallback without reapplying skip-worktree is also semantically wrong: excluded paths become live H entries and can look like deletions. Open #11665 only retries transient exit failures (not timeouts) and would not fix this sparse-path exit.

    Related, not duplicates. #3646 is the umbrella large-repo 30s add -A timeout. #10792 is the partial reuse optimization that correctly opts out of this shape. #11665 is exit retries only. No open PR teaches capture about sparse-checkout.

    Suggested fix (in GitVcsDriver.checkpoints.captureCheckpoint). Detect sparse-checkout and keep / reuse skip-worktree (those S flags are the cone). Stop dumping a 40 MiB ls-files -v to decide reuse. Make git add sparse-safe (in-cone pathspecs or reapply sparse on the temp index; do not fail capture on outside-cone untracked warnings). If fallback still runs, reapply skip-worktree before add. Tests: cone-mode + sparse index + outside-cone untracked + in-cone edit; capture under 30s; workspace index unchanged; tree has the in-cone change and not 211k excluded deletions.

    Workaround. None in-app. Removing the six outside-cone files would only expose the 30s fallback timeout; reuse would still be rejected. Disabling sparse-checkout is not realistic here. Do not treat “update to 0.0.42”, “bump the timeout”, or “land #11665” as closing this.

  2. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 16, 2026
  3. vedprakash2302 commented on Sep 16, 2026

    @vedprakash2302
    ContributorAuthor

    Opened #12154 to address this. It preserves sparse-checkout index reuse, removes the listing-size cutoff from the reuse decision, and handles files outside the sparse rules. Synthetic regression tests and benchmarks are in the PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions