Repository navigation
[Bug]: Checkpoint capture fails with exit 128 on every turn for repos on SMB shares (macOS): core.fsyncMethod=fsync → F_FULLFSYNC not supported #14402
Description
Activity
Triage
I confirmed this is a real server-side regression on current
main(c18e5ea, nightlyv0.0.45-nightly.20260930.2481) as well as the0.0.44you reported. It isn't a duplicate of #10905 (open issue) or #4380 (open issue). Updating won't help yet: the flags that fail came in with merged PR #10944, which shipped inv0.0.43,v0.0.44, and is still onmain.What's happening
In
apps/server/src/vcs/GitVcsDriver.ts,captureCheckpointrunsadd,write-tree,commit-tree, andupdate-refthroughdurableWrite, which adds:-c core.fsync=objects,reference -c core.fsyncMethod=fsyncA test ("flushes checkpoint objects and refs to disk before publishing them") pins those flags. Because
-con the command line beats repo config, settingcore.fsyncMethodlocally can't work around it.On macOS,
core.fsyncMethod=fsyncmakes Git callfcntl(F_FULLFSYNC). The macOS default,writeout-only, uses plainfsync(2)instead. SMB (smbfs) returnsENOTSUPforF_FULLFSYNC, so Git dies with the exact line you saw:fatal: fsync error on '.git/objects/xx/tmp_obj_*': Operation not supportedGit doesn't remove the temp object when it dies, so each failed
git addleaves atmp_obj_*file behind untilgit prune.batchdoesn't help because it still does a hardware flush; onlywriteout-onlytakes the plainfsyncpath, which matches your table. Other macOS volumes that rejectF_FULLFSYNC(NFS, AFP) should hit the same thing. Linux isn't affected, sincefsyncthere is justfsync(2).One detail if you keep the PATH wrapper: commands run as
git -C <repo>from the server's own working directory, so the wrapper needs to look at-C, not the cwd.Why the error is opaque, and why it looks like retries
VcsProcessExitError.fromProcessExitkeeps only the stderr length and replaces the text with "Process exited with a non-zero status.", which is what shows up in the activity log. That's the same problem as #4380, and it's why thefsync errorline never reaches the UI.These aren't retries. The checkpoint retry from merged PR #11665 only kicks in for a lock-file
File existsor a missing-pathNo such file or directory. The extratmp_obj_*files come from separate captures in each turn (the pre-turn baseline, possibly again onturn.started, and the completion capture), and each one dies insidegit add.Fix direction
We should keep
durableWriteon APFS, since #10944 exists to stop a crash from leaving a 0-byterefs/t3/**file that breaks later fetch and push (#10905). The better fix is to recognize this specificENOTSUPfsync error inVcsProcessbefore stderr is discarded, then retry that capture withcore.fsyncMethod=writeout-only. On SMB that's the strongest flush the volume actually supports, so durability isn't really lost. Retrying every exit 128 blindly would hide real disk failures, so the check should be narrow. Checking the filesystem type up front (smbfs,nfs,afpfs,webdav) could skip the doomed first attempt, and theENOTSUPretry covers anything that list misses. Once the failinggit addstops, thetmp_obj_*leak stops too.Until then, your PATH wrapper is a valid workaround; there's no config-only one. Thanks for the really thorough measurements.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 30, 2026 Fix is in #14411. The first attempt still uses full fsync. Only the exact
fsync error on …: Operation not supportedfailure retries that command once withwriteout-only.
Before submitting
Area
apps/server
Steps to reproduce
smbfs) that contains a git repository, e.g./Volumes/share.Minimal repro without T3 — this is exactly what
captureCheckpointdoes viadurableWrite:Expected behavior
Checkpoints are captured, as they were before
core.fsyncwas added to the capture path. If the filesystem cannot do a full fsync, T3 should fall back (e.g. retry withoutdurableWrite, or usecore.fsyncMethod=writeout-only) instead of failing the checkpoint.Actual behavior
Every turn that changes something shows
VCS process failed in GitVcsDriver.checkpoints.captureCheckpoint: git (/Volumes/share) exited with 128 - Process exited with a non-zero status.No
refs/t3/checkpoints/**are written, so undo/diff for the thread is unavailable. Each failed attempt (including the retries) leaves an orphaned.git/objects/xx/tmp_obj_*file behind; I found dozens of them after a day.Root cause:
captureCheckpointpasses-c core.fsync=objects,reference -c core.fsyncMethod=fsynctoadd,write-tree,commit-treeandupdate-ref. On macOS,fsyncMethod=fsyncmeansfcntl(F_FULLFSYNC), which the macOS SMB client does not support (ENOTSUP), and git aborts with exit 128.Measured on the same SMB share:
core.fsyncMethodfsync(what T3 passes)batchwriteout-onlyfsyncon local APFSRepo-level config cannot work around it, because
-con the command line always wins.Diagnosing this was hard because the git stderr (
fsync error … Operation not supported) is not surfaced anywhere, only the exit code (same class as #4380). Also note that the checkpoint git commands run withgit -C <repo>from the service's working directory, not with the repo as cwd.Impact
Major degradation or frequent failure
Version or commit
0.0.44
Environment
macOS 27.0.1 (Apple Silicon), git 2.56.0 (Homebrew; Apple Git 2.54.0 behaves the same), repository on an SMB share mounted via Finder (
smbfs), Claude providerLogs or stack traces
Workaround
A
gitwrapper placed before the real git in the service'sPATH(LaunchAgentEnvironmentVariables:PATH) that rewrites-c core.fsyncMethod=fsyncto-c core.fsyncMethod=writeout-onlywhen the target repo (cwd or-Cargument) is under/Volumes/. With that, checkpoints are captured again. Note thatlaunchctl kickstart -kdoes not pick up a changed PATH in the plist;bootout+bootstrapis needed.Possible fixes, in order of preference:
ENOTSUP/fsync error, retry the capture withoutdurableWrite(or withcore.fsyncMethod=writeout-only) instead of failing.statfs→smbfs,nfs,afpfs,webdav) and skipdurableWritethere.tmp_obj_*left behind by a failed capture.Related: #10905 (the reason
durableWriteexists), #4380 (opaque git errors).