Skip to content

[Bug]: Checkpoint capture fails with exit 128 on every turn for repos on SMB shares (macOS): core.fsyncMethod=fsync → F_FULLFSYNC not supported #14402

Description

@derstrassi

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. On macOS, mount an SMB share (smbfs) that contains a git repository, e.g. /Volumes/share.
  2. Open that folder as a project in T3 Code.
  3. Let an agent change any file (anything that makes the checkpoint write at least one new object).

Minimal repro without T3 — this is exactly what captureCheckpoint does via durableWrite:

cd /Volumes/share            # git repo on an smbfs mount
echo test | git -c core.fsync=objects,reference -c core.fsyncMethod=fsync hash-object -w --stdin
# fatal: fsync error on '.git/objects/a5/tmp_obj_Pseg66': Operation not supported   (exit 128)

Expected behavior

Checkpoints are captured, as they were before core.fsync was added to the capture path. If the filesystem cannot do a full fsync, T3 should fall back (e.g. retry without durableWrite, or use core.fsyncMethod=writeout-only) instead of failing the checkpoint.

Actual behavior

Every turn that changes something shows
VCS process failed in GitVcsDriver.checkpoints.captureCheckpoint: git (/Volumes/share) exited with 128 - Process exited with a non-zero status.
No refs/t3/checkpoints/** are written, so undo/diff for the thread is unavailable. Each failed attempt (including the retries) leaves an orphaned .git/objects/xx/tmp_obj_* file behind; I found dozens of them after a day.

Root cause: captureCheckpoint passes -c core.fsync=objects,reference -c core.fsyncMethod=fsync to add, write-tree, commit-tree and update-ref. On macOS, fsyncMethod=fsync means fcntl(F_FULLFSYNC), which the macOS SMB client does not support (ENOTSUP), and git aborts with exit 128.

Measured on the same SMB share:

core.fsyncMethod result
fsync (what T3 passes) fails, exit 128
batch fails, exit 128
writeout-only works
fsync on local APFS works

Repo-level config cannot work around it, because -c on the command line always wins.

Diagnosing this was hard because the git stderr (fsync error … Operation not supported) is not surfaced anywhere, only the exit code (same class as #4380). Also note that the checkpoint git commands run with git -C <repo> from the service's working directory, not with the repo as cwd.

Impact

Major degradation or frequent failure

Version or commit

0.0.44

Environment

macOS 27.0.1 (Apple Silicon), git 2.56.0 (Homebrew; Apple Git 2.54.0 behaves the same), repository on an SMB share mounted via Finder (smbfs), Claude provider

Logs or stack traces

VcsProcessExitError: VCS process failed in GitVcsDriver.checkpoints.captureCheckpoint: git (/Volumes/share) exited with 128 - Process exited with a non-zero status.
    at VcsProcessExitError.fromProcessExit
    at GitVcsDriver.checkpoints.captureCheckpoint
    at captureCheckpoint
    at ensurePreTurnBaselineFromTurnStart
    at processRuntimeEvent

# the underlying git error, captured via a wrapper:
fatal: fsync error on '.git/objects/4f/tmp_obj_muChk3': Operation not supported

Workaround

A git wrapper placed before the real git in the service's PATH (LaunchAgent EnvironmentVariables:PATH) that rewrites -c core.fsyncMethod=fsync to -c core.fsyncMethod=writeout-only when the target repo (cwd or -C argument) is under /Volumes/. With that, checkpoints are captured again. Note that launchctl kickstart -k does not pick up a changed PATH in the plist; bootout + bootstrap is needed.

Possible fixes, in order of preference:

  • On ENOTSUP/fsync error, retry the capture without durableWrite (or with core.fsyncMethod=writeout-only) instead of failing.
  • Or detect network filesystems (statfs → smbfs, nfs, afpfs, webdav) and skip durableWrite there.
  • Clean up tmp_obj_* left behind by a failed capture.

Related: #10905 (the reason durableWrite exists), #4380 (opaque git errors).

Activity

  1. juliusmarminge commented on Sep 30, 2026

    @juliusmarminge
    Member

    Triage

    I confirmed this is a real server-side regression on current main (c18e5ea, nightly v0.0.45-nightly.20260930.2481) as well as the 0.0.44 you reported. It isn't a duplicate of #10905 (open issue) or #4380 (open issue). Updating won't help yet: the flags that fail came in with merged PR #10944, which shipped in v0.0.43, v0.0.44, and is still on main.

    What's happening

    In apps/server/src/vcs/GitVcsDriver.ts, captureCheckpoint runs add, write-tree, commit-tree, and update-ref through durableWrite, which adds:

    -c core.fsync=objects,reference -c core.fsyncMethod=fsync
    

    A test ("flushes checkpoint objects and refs to disk before publishing them") pins those flags. Because -c on the command line beats repo config, setting core.fsyncMethod locally can't work around it.

    On macOS, core.fsyncMethod=fsync makes Git call fcntl(F_FULLFSYNC). The macOS default, writeout-only, uses plain fsync(2) instead. SMB (smbfs) returns ENOTSUP for F_FULLFSYNC, so Git dies with the exact line you saw:

    fatal: fsync error on '.git/objects/xx/tmp_obj_*': Operation not supported
    

    Git doesn't remove the temp object when it dies, so each failed git add leaves a tmp_obj_* file behind until git prune. batch doesn't help because it still does a hardware flush; only writeout-only takes the plain fsync path, which matches your table. Other macOS volumes that reject F_FULLFSYNC (NFS, AFP) should hit the same thing. Linux isn't affected, since fsync there is just fsync(2).

    One detail if you keep the PATH wrapper: commands run as git -C <repo> from the server's own working directory, so the wrapper needs to look at -C, not the cwd.

    Why the error is opaque, and why it looks like retries

    VcsProcessExitError.fromProcessExit keeps only the stderr length and replaces the text with "Process exited with a non-zero status.", which is what shows up in the activity log. That's the same problem as #4380, and it's why the fsync error line never reaches the UI.

    These aren't retries. The checkpoint retry from merged PR #11665 only kicks in for a lock-file File exists or a missing-path No such file or directory. The extra tmp_obj_* files come from separate captures in each turn (the pre-turn baseline, possibly again on turn.started, and the completion capture), and each one dies inside git add.

    Fix direction

    We should keep durableWrite on APFS, since #10944 exists to stop a crash from leaving a 0-byte refs/t3/** file that breaks later fetch and push (#10905). The better fix is to recognize this specific ENOTSUP fsync error in VcsProcess before stderr is discarded, then retry that capture with core.fsyncMethod=writeout-only. On SMB that's the strongest flush the volume actually supports, so durability isn't really lost. Retrying every exit 128 blindly would hide real disk failures, so the check should be narrow. Checking the filesystem type up front (smbfs, nfs, afpfs, webdav) could skip the doomed first attempt, and the ENOTSUP retry covers anything that list misses. Once the failing git add stops, the tmp_obj_* leak stops too.

    Until then, your PATH wrapper is a valid workaround; there's no config-only one. Thanks for the really thorough measurements.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 30, 2026
  3. mdshzb04 commented on Sep 30, 2026

    @mdshzb04
    Contributor

    Fix is in #14411. The first attempt still uses full fsync. Only the exact fsync error on …: Operation not supported failure retries that command once with writeout-only.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions