Repository navigation
fix: keep producer LaunchAgents registered and running (#142) - #143
Merged
Merged
Conversation
The CLI install/uninstall subcommands and Ember.app's SMAppService registration share the com.ember.heartbeat / com.ember.codex labels, and the CLI ran a blind `launchctl bootout gui/<uid>/<label>`. launchd drops a booted-out submitted job for good while Background Items keeps showing it enabled, so the daemon stays dead until the next login. That is what happened on m4 (#142): the unified log shows "bootout initiated by: launchctl <- ember-claude-pr <- go", i.e. the uninstall test calling uninstallPlist with the real uid during `go test`, at 2026-09-26 00:11:51. Every `go test ./...` on a Mac with the app's reporting on killed the heartbeat daemon. Now the CLI checks `launchctl print` for managed_by ServiceManagement: install refuses, uninstall leaves the app's job alone, and the test injects a fake launchctl instead of touching launchd.
… lost them Two gaps let an enabled producer stay dead (#142): - The update reconcile keyed on CFBundleVersion, which is always "1", so no update (release or a local rebuild ditto'd over /Applications) ever re-registered the agents. The key is now marketing version + build + a SHA-256 prefix of each bundled helper and plist, so any change to what launchd runs triggers one reconcile. - SMAppService.status reads the Background Items database, which stays .enabled after launchd drops the job (booted out, or unspawnable). On every launch each enabled agent is now probed with `launchctl print gui/<uid>/<label>` and re-registered if it's missing. The decision is a pure reconcileReason(registration:loaded:bundleChanged:) and reconcile(bundleChanged:) reports per-agent outcomes instead of stopping at the first error; the fingerprint is recorded only when every re-registration succeeded. AgentState gains .notRunning for the UI.
An agent that's on but has no launchd job used to read "On" while nothing reported. Its row now says "Not running" with a Repair button (re-registers it via ProducerInstallService.repairAll), and the section footer explains the state. Reporting stays on: the toggle reflects what the user chose.
CFBundleVersion stayed "1" for every release, so anything keyed on it (the producer re-register check, #142) never saw an update. The app no longer relies on it alone, but a build number that never moves is still wrong; release.sh now increments it alongside MARKETING_VERSION.
Why an enabled SMAppService agent can be dead, how the app now detects and repairs it, and why the CLI must leave app-registered labels alone (#142).
tarakanof
commented
Sep 26, 2026
tarakanof
left a comment
Owner
Author
There was a problem hiding this comment.
Review: #143 (fixes #142)
Verdict: fix #1 first (it's cheap), the rest can follow. The root cause is right and the test hole is closed. I verified the following:
- Root cause confirmed. On
main,uninstallPlist(home, uid)callsexec.Command("launchctl", "bootout", "gui/<uid>/com.ember.heartbeat"), andTestUninstall_StripsHooksAndRemovesPlistpassesos.Getuid(). That's a real bootout of the app's label. - No test reaches launchd any more. Grepped all Go packages (claude + codex producer, internal/producer): mutating launchctl calls are only reachable via
install(),runUninstall(),reloadLaunchAgentanduninstallPlist, and the only test caller (uninstallPlist) now gets a fake. Rango test ./... -race -count=1with alaunchctlshim first on PATH that logs and fails: 0 calls. Thecom.ember.codexpid was unchanged afterwards. Swift: onlyAppEnvironmentbuildsRealSMAppService/ProcessCommandRunner, and every test uses fakes. launchctl printformat checked on this Mac (macOS 27): the SMAppService job showsmanaged_by = com.apple.xpc.ServiceManagement,type = Submitted,path = (submitted by smd.339). A missing job exits 113.- Fingerprint cost: the two helpers total about 39 MB, and
shasum -a 256takes about 0.12 s. It runs inTask.detached(.utility), so that's fine. A missing helper hashes as"missing", which is stable. release.sh: simulated the sed onmacos/project.yml. OnlyCURRENT_PROJECT_VERSION: "1"changes, to"2". The numeric read fails loudly, andproject.ymlis what gets committed. OK.swift test --package-path macos: 470 pass. Unsigned Debugxcodebuildsucceeds with no Swift warnings.go vetis clean. CI is green.
Should-fix
- Repair/reconcile may never reach
register()in exactly the #142 state.reconciledoestry sm.unregister(...)thentry sm.register(...). For a job launchd has dropped while Background Items still says enabled, I couldn't verify on-device thatSMAppService.unregister()succeeds. If it throws,register()is skipped and the error is recorded. The fingerprint then stays unrecorded, launch retries forever, and Repair always fails. Cheap hardening for.notRunning:try? unregister()thentry register(), or register first. Add a test with a throwingunregister. Please also check this in the on-device verification. managed_bydetection fails open.appManagedis an exact substring match on undocumentedlaunchctl printoutput. If Apple renames or reformats the key,CheckNotAppManagedreturns nil and the CLI boots out the app's job again, silently. Safer: boot out only on positive identification of the CLI job (path = ~/Library/LaunchAgents/<label>.plist). Or at least also treattype = Submitted/(submitted by smdas app-managed.
Nits
isLoadedtreats any non-zero exit as "not loaded", but only 113 means "Could not find service". Any other failure re-registers on every launch and every Settings refresh. That contradicts the doc comment's "a broken probe never churns registrations". SuggestexitCode != 113means loaded.- In both producers,
install()runsconfigure()(writes~/.claude/settings.json/producer.env) before theCheckNotAppManagedrefusal. It then errors after already editing config. Move the check to the top. - In the #142 state (app registration enabled but job dropped),
printfails, so CLIinstallbootstraps its own plist under the app's label. The app then reports.onfor a CLI-owned job. This is an edge case; maybe worth a RUNBOOK line. - The
doctorhint for a missing heartbeat is stilllaunchctl bootstrap gui/<uid> <cli plist>. On an app-managed setup that plist doesn't exist, and the right action is Settings › Agents › Repair. This is optional and adjacent to the PR's scope. - The "record fingerprint only if every outcome succeeded" logic lives in the app target (
AppEnvironment.reconcileProducers), so no test covers it. It's small, but it could be a pure EmberKit helper.
Repair UX looks fine: a per-row "Not running" state with a Repair button, the toggle stays on, a footer explains it, and the strings were added to the catalog.
…ring Review of #143: the managed_by match failed open, so a reformatted launchctl print would let the CLI boot out the app's job again. The CLI now boots out only a job whose path is its own plist; SMAppService markers, other paths, unparseable output or a launchctl failure count as someone else's. install checks before configure() edits settings.json/producer.env, and refuses when launchd holds the app's enabled override for the label even though the job was dropped (the #142 state). doctor points app setups at Settings › Agents › Repair.
Review of #143: unregister may throw for a registration launchd already dropped, which skipped register and left Repair and the launch retry failing forever; it's now best-effort. Only exit 113 / 'Could not find service' means not loaded; other launchctl failures count as loaded and are logged once, so a broken probe can't churn registrations. The record-the-fingerprint decision moves into EmberKit (shouldRecordFingerprint) with tests.
This was referenced Sep 26, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #142
Root cause
The unified log on m4 shows who removed the heartbeat job:
ember-claude-prundergois theember-claude-producer.testbinary.TestUninstall_StripsHooksAndRemovesPlistcallsuninstallPlist(tmp, os.Getuid()), which ran a reallaunchctl bootout gui/501/com.ember.heartbeat. The CLI and Ember.app's SMAppService registration use the same label, so everygo test ./...on a Mac with reporting on killed the app's heartbeat daemon. launchd drops a booted-out job for good, but Background Items still shows it enabled until the next login.SMAppService.statusreads the same database and returned.enabled, so nothing noticed. Codex survived because no test boots it out.The CFBundleVersion "1" key from the issue is a second bug: no update ever triggered a re-register. It didn't cause this outage (the job was dropped at 00:11, about 14 hours before the app was replaced), but with ad-hoc builds it would. The codex job's
LWCRpins the helper cdhash, and a rebuilt helper gets a new one.Changes
internal/producer/launchagent.go. The CLIinstallrefuses, anduninstallskips the bootout, whenlaunchctl printshowsmanaged_by = com.apple.xpc.ServiceManagement. Only CLI-loaded jobs get booted out. launchctl is injected, and the uninstall test uses a fake, so tests never touch launchd again.bundleFingerprint). The stored"1"never matches, so the first launch of this build reconciles once. At every launch, each enabled agent is probed withlaunchctl print gui/<uid>/<label>and re-registered if launchd has no job. The decision logic is the purereconcileReason(registration:loaded:bundleChanged:).reconcile(bundleChanged:)returns an outcome per agent, and the fingerprint is recorded only when every re-register succeeded.AgentState.notRunningas "Not running" with a Repair button (repairAll()), plus a footer note. The toggle stays on.release.shnow incrementsCURRENT_PROJECT_VERSIONtoo.Tests
go test ./... -race,go vet, gofmt (1.26): pass.launchctl print gui/501/com.ember.codexstill shows the same pid after the run, so the tests no longer touch launchd.swift test --package-path macos: 470 pass. NewProducerReconcileTestscover the reason matrix, the probe args, a failing probe counting as loaded, notRunning vs toggle, snapshotneedsRepair, reconcile (unloaded only / all after an update / healthy no-op / per-agent failures), repairAll, and the fingerprint (a helper change with the same version, a version or build change, legacy"1"). There is also a model repair test.xcodebuild ... CODE_SIGNING_ALLOWED=NO build: succeeds with no new warnings.scripts/strings.sh syncadded 4 keys, andcheckpasses.Needs on-device verification (not done here)
After installing this build on m4, the first launch should re-register both agents (reason=bundleChanged in
log show --predicate 'subsystem == "com.ember.Ember"'), andlaunchctl print gui/501/com.ember.heartbeatshould find the job. I did not change any launchd state on this Mac.