Repository navigation
runtime-host: intermittent 'stopped responding during startup' after a Windows upgrade over an existing profile #3279
Description
Activity
Claiming the verifier half (capturing the Host's own stderr/log into the lifecycle failure output); the reproduction and any Host-side fix I will attempt after the open release-verifier PRs land.
New occurrence with app-side evidence, and a scope widening: this reproduces on a fresh install with a fully isolated profile, not only after an upgrade over an existing profile.
Run 32380815350 (
release-windows-check, head 83ba1fd), step "Verify the Windows release", stage "smoking the packaged renderer" — the very first packaged launch in the job: freshly installed candidate, isolated--user-data-dir, isolated HOME/APPDATA/LOCALAPPDATA. The harness (with the newly added stderr capture from #3327) reports:Error: Packaged Maka renderer did not expose CDP within 90 seconds: The operation was aborted due to timeout. [runtime-host] fatal: Error: Runtime Host stopped responding during startup [startup] fatal: Error: Runtime Host stopped responding during startupMapping to source:
host_unresponsivefromconnect-or-spawn.ts— the election loop saw an endpoint but it never became responsive withinDEFAULT_ELECTION_DEADLINE_MS = 45_000, thensawUnresponsiveEndpointselectshost_unresponsive(connect-or-spawn.ts:340). With an isolated fresh profile there is no pre-existing endpoint to inherit, so the unresponsive endpoint was created by this same app instance's own spawn on a heavily loaded runner.Two implications for whoever picks this up:
- The earlier bare "did not expose CDP / fetch failed" failures on the Windows lanes (at least four across feat(windows-sandbox): implement production-identity readiness probe #3161/feat(release): verify Windows automatic updates end to end #3240/feat(win): Abort-path installer rollback with backup retention #3265 heads) are plausibly this same product-level startup failure — the harness could not see app stderr until now, so they were attributed to CDP attach. The capture is in place going forward, so future occurrences will self-identify.
- The 45s election deadline plus a wedged named-pipe endpoint on a slow runner is enough to kill startup outright (the app exits rather than retrying the spawn). Whether the right fix is a longer deadline, a respawn-on-unresponsive retry, or endpoint health hygiene is a product decision — but the failure is user-visible (app fails to launch), not just CI noise.
Evidence grade: run log + source read (
connect-or-spawn.ts:37,340,startup-error.ts:52); not locally reproduced.- added a commit that references this issue
on Aug 20, 2026 Status after #3382 and the latest #3265 Windows run:
- fix(desktop): drain untracked Runtime Host before update #3382/windows: auto-update silently fails when a Runtime Host process survives the app quit (installer cannot clear it) #3340 is a distinct update-handoff process-lifetime bug and is now closed; it should not be treated as resolving this startup issue.
- Release Windows check 32477989412 on feat(win): Abort-path installer rollback with backup retention #3265 exact head
68520adccompleted cleanly in 23m21s. It launched packaged Maka repeatedly across clean release smoke, pinned-version upgrade, automatic update/relaunch, rollback, recovery reruns, and uninstall. Nohost_unresponsiveoccurred in that run. - This is one useful non-occurrence, not resolution evidence. The fresh-isolated-profile failure already recorded here means predecessor state is not a necessary trigger.
Recommended next product investigation:
- Run a repeated fresh-profile startup stress loop under CPU/disk pressure while capturing the execution candidate/Host child's stderr, exit status, endpoint publication time, first successful pipe response, and election deadline.
- Classify the failing cut before changing timeouts: child exited, child alive with a published-but-unresponsive pipe, or stale endpoint/registration observed by election.
- If retry is added, bind it to the exact manager-owned candidate/Host identity and prove cleanup before respawn; do not turn the 45s deadline into an unbounded wait or hide the fault with verifier retries.
The verifier already preserves app-side fatal output, so future occurrences should stay attributable to this issue.
中文:最新 #3265 Windows 全链路多次启动均通过,本轮未复现
host_unresponsive,但单次未复现不能关闭问题;全新 profile 已证明旧状态不是必要条件。下一步应做带 Host 子进程 stderr/退出码/pipe 时间线的压力复现,先定位 child exit、活着但 endpoint 无响应、还是 election 观察到 stale endpoint,再决定是否做有身份约束的单次 respawn。@Joob1n I am starting a narrow product-side diagnostic slice for #3279 from current main, avoiding the verifier surface you previously claimed.
Scope:
- preserve candidate/Host child stderr + exit/endpoint/election timing as bounded structured startup evidence;
- add a repeated fresh-profile startup stress entry point for Windows;
- do not increase the election deadline or add retry/respawn until the failing cut is classified;
- no changes to the release verifier's existing CDP/process-probe ownership unless a product diagnostic needs a minimal read seam.
I will post the exact branch/PR and evidence before proposing a behavior fix. Please flag any active overlap and I will rebase the ownership boundary rather than duplicate work.
@liugddx No overlap — go ahead, and thank you for asking rather than assuming.
The verifier half I claimed on 2026-08-20 was capturing the Host's own stderr into the lifecycle failure output. That is already done, and you did it: #3327 landed it, and your comment above is reading the evidence it produces ("with the newly added stderr capture from #3327"). There is nothing left of my claim to protect.
The product half I said I would attempt after the release-verifier PRs landed. Those have landed and I did not start it, so it is yours with nothing to rebase around. I have unassigned myself so the issue does not read as covered when it is not.
Your scope looks right to me, and one part of it is the part I would have gotten wrong: not increasing the election deadline or adding retry/respawn until the failing cut is classified. A timeout that is widened before the cause is known stops reproducing without being fixed, and the evidence goes with it.
One observation, offered as context and not as evidence — it is from macOS and this issue is Windows. While measuring something unrelated on my own machine,
codesign --timestampround trips varied between 285 ms and 2019 ms on the same binary, same machine, minutes apart. Nothing about that is specific to codesign; it is what an unremarkable network round trip looks like when the runner is not the constant it appears to be. If the fresh-profile stress loop can vary CPU/disk pressure between iterations rather than holding it steady, a startup deadline that only fails under a particular ratio of contention would surface faster than one measured at a fixed load. Discard it if the failing cut turns out to be deterministic.I will stay out of the way. If a product diagnostic ends up needing a read seam into the verifier's CDP or process-probe surface, ping me — I know that code and can review the seam quickly rather than have you work around it.
Correction to my last comment: I said I had unassigned myself, and I could not. Both paths are refused for an outside contributor —
ReplaceActorsForAssignableon the GraphQL side, and the REST assignees endpoint the same way.So the assignment still reads
Joob1nand that is now wrong: nothing on this issue is mine, and the half I did claim was delivered by #3327. @liugddx, if you have the permission, please take the assignment — or any committer can drop me from it. Until then the field says this issue is covered when it is not, which is the opposite of what I want it to say.Product-side diagnostic PR: #3454 (
b2315f21eec3076592d86643d492c6e7f5e70183).Scope correction to my earlier kickoff comment: I wrote "child stderr + exit/endpoint/election timing," which was stronger than the narrow implementation and should not be read as the PR contract. The detached Candidate still uses
stdio: ignore; #3454 does not take ownership of child stderr or claim exact endpoint publication/phase timestamps. Adding pipes here would change detached-process I/O and lifecycle before the failing cut is known.The actual bounded evidence is:
- exact Candidate PID and startup-attempt identity;
- running/exited/unknown state with exit code/signal when observed;
- final election elapsed time and true process-attempt count (including coalesced launches);
- endpoint-connected observation plus fixed election counters;
- the last safe registration state/PID/lifecycle/generation summary.
No root path or endpoint value is added to the startup error. The PR also adds a fresh-root Windows stress runner with per-iteration NDJSON and exact Candidate settlement. It does not change the 45s deadline, retry/respawn, backoff, endpoint behavior, or failure classification.
On the exact PR head, a fresh-root Windows run connected 1/1 in 5.642s; the exact Candidate exited cleanly and the new root was removed. The PR is draft pending core CI and
Release Windows checkon the merged result.CI follow-up for #3454 exact head
b2315f21eec3076592d86643d492c6e7f5e70183:- core CI 32537529458 passed in 13m01s, including the full Runtime Host suite and Desktop E2E;
- Dependency audit 32537529491 passed in 41s;
- Release Windows check 32537529459 passed in 14m05s, including packaged smoke, pinned upgrade/uninstall, and automatic update end to end.
The PR is now ready for review. It remains a diagnostic slice, not a claim that #3279's startup behavior is fixed.
I also attempted the assignment correction requested above, but GitHub rejected it with
ReplaceActorsForAssignable; the issue still lists @Joob1n despite the explicit ownership handoff. A committer with assignment permission will need to remove or replace that stale field.#3454 produced a classified packaged recurrence on exact head
8ab1ce2cd78f4f5fe7eadc5bd981b1aa4f65e00b.Release Windows check 32567384310, final "Verify automatic update end to end" startup:
{ "deadlineMs": 45000, "elapsedMs": 45174, "candidateLaunches": 69, "sawEndpointConnected": true, "observations": { "notRegistered": 73, "connectFailed": 0, "handshakeFailed": 0, "connected": 1, "readyWaitFailed": 1, "deadlineElapsed": 0 }, "lastRegistration": { "pid": 8268, "state": "recovering", "lifecycleMode": "ephemeral", "generation": "0.1.11" }, "latestCandidate": { "pid": 9812, "startupAttemptId": "8edc95c0-e691-46c5-93ca-eefe5b837cc6", "state": "running" } }Classification:
- not a pre-existing stale endpoint: this was a fresh automatic-update fixture and the endpoint was reached;
- not a Candidate that exited before publication: the latest exact attempt was still running;
- one Host accepted a connection but did not become ready before the election deadline; its last registration was
recovering; - before registration stabilized, the same election made 69 distinct process attempts while observing
not_registered73 times.
This gives a concrete behavior-fix path: keep at most one exact live Candidate attempt per election until that attempt exits or a registration/Host appears, then decide whether another launch is warranted. The fix must retain the existing bounded deadline and exact-attempt cleanup; it should not be implemented as a verifier retry or a wider timeout.
I am not adding that behavior change to #3454 because the PR is intentionally diagnostic-only. Core CI and audit on the exact head passed; the Windows lane passed packaging, packaged smoke, and pinned upgrade/uninstall before this product-level recurrence.
Behavior fix PR: #3512 (
814d95f663c74fb6bd804de5eed3ae37d2eae482).It addresses the concrete convergence failure captured by #3454 without expanding the diagnostic PR:
- one election keeps at most one Candidate process in flight;
- the gate opens only after that process's exact
exitedpromise settles; - spawn rejection remains retryable because no process was created;
- an unknown/rejected exit signal fails closed for that election;
- the production 45s deadline, readiness polling, failure classification, endpoint behavior, and cleanup ownership are unchanged;
- cross-client/global coalescing and full LaunchLease architecture remain out of scope.
Direct evidence includes pending/exit/spawn-reject table paths, a real owned Candidate exit followed by exactly one production successor, and a real Electron Client reporting exactly one Candidate PID. The two startup source files now trigger
Release Windows check, so packaged automatic-update evidence is required before the draft is opened for review.#3454 stays frozen until #3512 reaches a final disposition. If #3512 lands, #3454 can rebase, remove its Node-only stress harness, and close its separate counter-completeness P3 without mixing the behavior repair into the diagnostic PR.
#3512 exact-head CI is green:
- Core CI 32581049491 passed in 6m09s, including the full Runtime Host suite.
- Release Windows check 32581049489 passed in 12m16s, including packaged smoke, pinned upgrade/uninstall, and the automatic-update path that reproduced the 69-Candidate storm on test(runtime-host): attribute Windows startup stalls #3454.
The PR is now ready for human review. This is one successful exact-head packaged run, so it verifies the repaired path but does not justify broader global/cross-client coalescing claims. #3454 remains frozen until #3512 is reviewed and merged; after that it can rebase and be reduced to the permanent diagnostic surface.
Final status after #3512 and the trimmed #3454 rerun:
- fix(runtime-host): keep one candidate in flight #3512 is merged and owns the per-election Candidate single-flight behavior fix.
- test(runtime-host): attribute Windows startup stalls #3454 exact head
8a5a275926f7abef0109b496acd84dc86cae535ais now diagnostic-only: the exploratory Node stress harness was removed, endpoint phase evidence and total/other observation buckets remain, and the operator deadline guidance is preserved. - Core CI 32607957602 passed in 11m45s.
- Release Windows check 32607957610 passed in 11m38s, including packaged automatic-update end to end without the previous Candidate storm.
The issue now has a behavior fix merged and a permanent diagnostic PR ready for normal human review. It should remain open for any future recurrence, not be closed solely because one packaged run passed.
The behavior fix landed in #3512, which keeps at most one Candidate in flight per election and opens the gate only after that process's exit settles; that is the direct cause of the 69-launch storm captured here. #3454 landed the permanent startup diagnostics, and there has been no reported recurrence since.
Closing as fixed. Please reopen with a new classified capture if it happens again.
On the Windows upgrade-lifecycle check, the app installed by an upgrade intermittently fails to start: its own stderr reports
and the renderer never mounts, so the packaged smoke times out.
Observed: run 32324991998 on #3241, step
Exercise pinned-version upgrade and uninstall. The sequence that failed: the pinned 0.1.9 baseline installed and verified clean → upgraded to the current 0.1.11 build → the upgraded app's Runtime Host stopped responding during startup. The standalone smoke of the same 0.1.11 build passed minutes earlier in the same job, so the build starts fine against a fresh profile.What makes the upgrade path different: the lifecycle deliberately reuses one isolated HOME and user-data directory across both versions — that is the point of an upgrade test — so the 0.1.11 Runtime Host starts against whatever state 0.1.9's run left under that HOME (host socket/lock/config). Whether the hang is caused by that leftover state or is an unlucky cold-runner slowdown is exactly what the log cannot yet say: the fatal is a client-side startup timeout, and the Host process's own stderr is not captured in the verifier output.
Attribution note: this surfaced through #3241's new CDP diagnostics (the app wrote
DevToolsActivePort, the poll named the bound port, and the endpoint never answered — which is what pointed at the main process rather than the port plumbing). Some of the historicaldid not expose CDP within 30 secondsfailures in this step may share this cause, but that cannot be established retroactively from the old logs.Suggested next steps:
Not caused by #3241 (verifier-only changes); filed separately so the product-side question is trackable.