Skip to content

fix: keep producer LaunchAgents registered and running (#142) - #143

Merged
tarakanof merged 7 commits into
mainfrom
fix/142-agent-reregister
Sep 26, 2026
Merged

tarakanof merged 7 commits into
mainfrom
fix/142-agent-reregister

Conversation

@tarakanof

Copy link
Copy Markdown
Owner

Closes #142

Root cause

The unified log on m4 shows who removed the heartbeat job:

2026-09-26 00:11:51.698 launchd: [gui/501/com.ember.heartbeat [1632]] bootout initiated by:
  launchctl[40713] <- ember-claude-pr[39735] <- go[39463] <- zsh <- claude
2026-09-26 00:11:51.700 launchd: removing service: com.ember.heartbeat

ember-claude-pr under go is the ember-claude-producer.test binary. TestUninstall_StripsHooksAndRemovesPlist calls uninstallPlist(tmp, os.Getuid()), which ran a real launchctl bootout gui/501/com.ember.heartbeat. The CLI and Ember.app's SMAppService registration use the same label, so every go test ./... on a Mac with reporting on killed the app's heartbeat daemon. launchd drops a booted-out job for good, but Background Items still shows it enabled until the next login. SMAppService.status reads the same database and returned .enabled, so nothing noticed. Codex survived because no test boots it out.

The CFBundleVersion "1" key from the issue is a second bug: no update ever triggered a re-register. It didn't cause this outage (the job was dropped at 00:11, about 14 hours before the app was replaced), but with ad-hoc builds it would. The codex job's LWCR pins the helper cdhash, and a rebuilt helper gets a new one.

Changes

  • fix(producer): new internal/producer/launchagent.go. The CLI install refuses, and uninstall skips the bootout, when launchctl print shows managed_by = com.apple.xpc.ServiceManagement. Only CLI-loaded jobs get booted out. launchctl is injected, and the uninstall test uses a fake, so tests never touch launchd again.
  • fix(menu): update detection now keys on marketing version + build + a SHA-256 prefix of each bundled helper and plist (bundleFingerprint). The stored "1" never matches, so the first launch of this build reconciles once. At every launch, each enabled agent is probed with launchctl print gui/<uid>/<label> and re-registered if launchd has no job. The decision logic is the pure reconcileReason(registration:loaded:bundleChanged:). reconcile(bundleChanged:) returns an outcome per agent, and the fingerprint is recorded only when every re-register succeeded.
  • feat(menu): Settings › Agents shows a new AgentState.notRunning as "Not running" with a Repair button (repairAll()), plus a footer note. The toggle stays on.
  • chore(release): release.sh now increments CURRENT_PROJECT_VERSION too.
  • docs: RUNBOOK and an ARCHITECTURE gotcha.

Tests

  • go test ./... -race, go vet, gofmt (1.26): pass. launchctl print gui/501/com.ember.codex still shows the same pid after the run, so the tests no longer touch launchd.
  • swift test --package-path macos: 470 pass. New ProducerReconcileTests cover the reason matrix, the probe args, a failing probe counting as loaded, notRunning vs toggle, snapshot needsRepair, reconcile (unloaded only / all after an update / healthy no-op / per-agent failures), repairAll, and the fingerprint (a helper change with the same version, a version or build change, legacy "1"). There is also a model repair test.
  • Unsigned xcodebuild ... CODE_SIGNING_ALLOWED=NO build: succeeds with no new warnings. scripts/strings.sh sync added 4 keys, and check passes.

Needs on-device verification (not done here)

After installing this build on m4, the first launch should re-register both agents (reason=bundleChanged in log show --predicate 'subsystem == "com.ember.Ember"'), and launchctl print gui/501/com.ember.heartbeat should find the job. I did not change any launchd state on this Mac.

The CLI install/uninstall subcommands and Ember.app's SMAppService
registration share the com.ember.heartbeat / com.ember.codex labels, and
the CLI ran a blind `launchctl bootout gui/<uid>/<label>`. launchd drops a
booted-out submitted job for good while Background Items keeps showing it
enabled, so the daemon stays dead until the next login.

That is what happened on m4 (#142): the unified log shows
"bootout initiated by: launchctl <- ember-claude-pr <- go", i.e. the
uninstall test calling uninstallPlist with the real uid during `go test`,
at 2026-09-26 00:11:51. Every `go test ./...` on a Mac with the app's
reporting on killed the heartbeat daemon.

Now the CLI checks `launchctl print` for managed_by ServiceManagement:
install refuses, uninstall leaves the app's job alone, and the test
injects a fake launchctl instead of touching launchd.
… lost them

Two gaps let an enabled producer stay dead (#142):

- The update reconcile keyed on CFBundleVersion, which is always "1", so
  no update (release or a local rebuild ditto'd over /Applications) ever
  re-registered the agents. The key is now marketing version + build +
  a SHA-256 prefix of each bundled helper and plist, so any change to
  what launchd runs triggers one reconcile.
- SMAppService.status reads the Background Items database, which stays
  .enabled after launchd drops the job (booted out, or unspawnable). On
  every launch each enabled agent is now probed with
  `launchctl print gui/<uid>/<label>` and re-registered if it's missing.

The decision is a pure reconcileReason(registration:loaded:bundleChanged:)
and reconcile(bundleChanged:) reports per-agent outcomes instead of
stopping at the first error; the fingerprint is recorded only when every
re-registration succeeded. AgentState gains .notRunning for the UI.
An agent that's on but has no launchd job used to read "On" while nothing
reported. Its row now says "Not running" with a Repair button (re-registers
it via ProducerInstallService.repairAll), and the section footer explains
the state. Reporting stays on: the toggle reflects what the user chose.
CFBundleVersion stayed "1" for every release, so anything keyed on it (the producer re-register check, #142) never saw an update. The app no longer relies on it alone, but a build number that never moves is still wrong; release.sh now increments it alongside MARKETING_VERSION.
Why an enabled SMAppService agent can be dead, how the app now detects and repairs it, and why the CLI must leave app-registered labels alone (#142).

@tarakanof tarakanof left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: #143 (fixes #142)

Verdict: fix #1 first (it's cheap), the rest can follow. The root cause is right and the test hole is closed. I verified the following:

  • Root cause confirmed. On main, uninstallPlist(home, uid) calls exec.Command("launchctl", "bootout", "gui/<uid>/com.ember.heartbeat"), and TestUninstall_StripsHooksAndRemovesPlist passes os.Getuid(). That's a real bootout of the app's label.
  • No test reaches launchd any more. Grepped all Go packages (claude + codex producer, internal/producer): mutating launchctl calls are only reachable via install(), runUninstall(), reloadLaunchAgent and uninstallPlist, and the only test caller (uninstallPlist) now gets a fake. Ran go test ./... -race -count=1 with a launchctl shim first on PATH that logs and fails: 0 calls. The com.ember.codex pid was unchanged afterwards. Swift: only AppEnvironment builds RealSMAppService/ProcessCommandRunner, and every test uses fakes.
  • launchctl print format checked on this Mac (macOS 27): the SMAppService job shows managed_by = com.apple.xpc.ServiceManagement, type = Submitted, path = (submitted by smd.339). A missing job exits 113.
  • Fingerprint cost: the two helpers total about 39 MB, and shasum -a 256 takes about 0.12 s. It runs in Task.detached(.utility), so that's fine. A missing helper hashes as "missing", which is stable.
  • release.sh: simulated the sed on macos/project.yml. Only CURRENT_PROJECT_VERSION: "1" changes, to "2". The numeric read fails loudly, and project.yml is what gets committed. OK.
  • swift test --package-path macos: 470 pass. Unsigned Debug xcodebuild succeeds with no Swift warnings. go vet is clean. CI is green.

Should-fix

  1. Repair/reconcile may never reach register() in exactly the #142 state. reconcile does try sm.unregister(...) then try sm.register(...). For a job launchd has dropped while Background Items still says enabled, I couldn't verify on-device that SMAppService.unregister() succeeds. If it throws, register() is skipped and the error is recorded. The fingerprint then stays unrecorded, launch retries forever, and Repair always fails. Cheap hardening for .notRunning: try? unregister() then try register(), or register first. Add a test with a throwing unregister. Please also check this in the on-device verification.
  2. managed_by detection fails open. appManaged is an exact substring match on undocumented launchctl print output. If Apple renames or reformats the key, CheckNotAppManaged returns nil and the CLI boots out the app's job again, silently. Safer: boot out only on positive identification of the CLI job (path = ~/Library/LaunchAgents/<label>.plist). Or at least also treat type = Submitted / (submitted by smd as app-managed.

Nits

  1. isLoaded treats any non-zero exit as "not loaded", but only 113 means "Could not find service". Any other failure re-registers on every launch and every Settings refresh. That contradicts the doc comment's "a broken probe never churns registrations". Suggest exitCode != 113 means loaded.
  2. In both producers, install() runs configure() (writes ~/.claude/settings.json / producer.env) before the CheckNotAppManaged refusal. It then errors after already editing config. Move the check to the top.
  3. In the #142 state (app registration enabled but job dropped), print fails, so CLI install bootstraps its own plist under the app's label. The app then reports .on for a CLI-owned job. This is an edge case; maybe worth a RUNBOOK line.
  4. The doctor hint for a missing heartbeat is still launchctl bootstrap gui/<uid> <cli plist>. On an app-managed setup that plist doesn't exist, and the right action is Settings › Agents › Repair. This is optional and adjacent to the PR's scope.
  5. The "record fingerprint only if every outcome succeeded" logic lives in the app target (AppEnvironment.reconcileProducers), so no test covers it. It's small, but it could be a pure EmberKit helper.

Repair UX looks fine: a per-row "Not running" state with a Repair button, the toggle stays on, a footer explains it, and the strings were added to the catalog.

…ring

Review of #143: the managed_by match failed open, so a reformatted launchctl print would let the CLI boot out the app's job again. The CLI now boots out only a job whose path is its own plist; SMAppService markers, other paths, unparseable output or a launchctl failure count as someone else's. install checks before configure() edits settings.json/producer.env, and refuses when launchd holds the app's enabled override for the label even though the job was dropped (the #142 state). doctor points app setups at Settings › Agents › Repair.
Review of #143: unregister may throw for a registration launchd already dropped, which skipped register and left Repair and the launch retry failing forever; it's now best-effort. Only exit 113 / 'Could not find service' means not loaded; other launchctl failures count as loaded and are logged once, so a broken probe can't churn registrations. The record-the-fingerprint decision moves into EmberKit (shouldRecordFingerprint) with tests.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(menu): producer LaunchAgents never re-registered after an app update

1 participant