Skip to content

runtime-host: intermittent 'stopped responding during startup' after a Windows upgrade over an existing profile #3279

Description

@Joob1n

On the Windows upgrade-lifecycle check, the app installed by an upgrade intermittently fails to start: its own stderr reports

[runtime-host] fatal: Error: Runtime Host stopped responding during startup
    at runtimeHostStartupError (…/app.asar/node_modules/@maka/runtime-host/dist/client/startup-error.js:25:20)
    at RuntimeHostDesktopManagerImpl.connect (…/dist/main/runtime-host-desktop-manager.js:317:19)
    at async RuntimeHostReconnectLifecycleImpl.start (…/reconnect-lifecycle.js:62:45)

and the renderer never mounts, so the packaged smoke times out.

Observed: run 32324991998 on #3241, step Exercise pinned-version upgrade and uninstall. The sequence that failed: the pinned 0.1.9 baseline installed and verified clean → upgraded to the current 0.1.11 build → the upgraded app's Runtime Host stopped responding during startup. The standalone smoke of the same 0.1.11 build passed minutes earlier in the same job, so the build starts fine against a fresh profile.

What makes the upgrade path different: the lifecycle deliberately reuses one isolated HOME and user-data directory across both versions — that is the point of an upgrade test — so the 0.1.11 Runtime Host starts against whatever state 0.1.9's run left under that HOME (host socket/lock/config). Whether the hang is caused by that leftover state or is an unlucky cold-runner slowdown is exactly what the log cannot yet say: the fatal is a client-side startup timeout, and the Host process's own stderr is not captured in the verifier output.

Attribution note: this surfaced through #3241's new CDP diagnostics (the app wrote DevToolsActivePort, the poll named the bound port, and the endpoint never answered — which is what pointed at the main process rather than the port plumbing). Some of the historical did not expose CDP within 30 seconds failures in this step may share this cause, but that cannot be established retroactively from the old logs.

Suggested next steps:

  1. Capture the Runtime Host child's stderr (or its log file) into the lifecycle verifier's failure output, so a startup hang is attributable to the Host's own last words rather than the client's timeout.
  2. Reproduce locally: install 0.1.9's profile state, then start a current build against it, in a loop.
  3. If leftover state is the trigger, the fix likely belongs in Host startup's handling of a predecessor's socket/lock remnants.

Not caused by #3241 (verifier-only changes); filed separately so the product-side question is trackable.

Activity

  1. Joob1n commented on Aug 20, 2026

    @Joob1n
    MemberAuthor

    Claiming the verifier half (capturing the Host's own stderr/log into the lifecycle failure output); the reproduction and any Host-side fix I will attempt after the open release-verifier PRs land.

  2. liugddx commented on Aug 20, 2026

    @liugddx
    Member

    New occurrence with app-side evidence, and a scope widening: this reproduces on a fresh install with a fully isolated profile, not only after an upgrade over an existing profile.

    Run 32380815350 (release-windows-check, head 83ba1fd), step "Verify the Windows release", stage "smoking the packaged renderer" — the very first packaged launch in the job: freshly installed candidate, isolated --user-data-dir, isolated HOME/APPDATA/LOCALAPPDATA. The harness (with the newly added stderr capture from #3327) reports:

    Error: Packaged Maka renderer did not expose CDP within 90 seconds: The operation was aborted due to timeout.
    [runtime-host] fatal: Error: Runtime Host stopped responding during startup
    [startup] fatal: Error: Runtime Host stopped responding during startup
    

    Mapping to source: host_unresponsive from connect-or-spawn.ts — the election loop saw an endpoint but it never became responsive within DEFAULT_ELECTION_DEADLINE_MS = 45_000, then sawUnresponsiveEndpoint selects host_unresponsive (connect-or-spawn.ts:340). With an isolated fresh profile there is no pre-existing endpoint to inherit, so the unresponsive endpoint was created by this same app instance's own spawn on a heavily loaded runner.

    Two implications for whoever picks this up:

    1. The earlier bare "did not expose CDP / fetch failed" failures on the Windows lanes (at least four across feat(windows-sandbox): implement production-identity readiness probe #3161/feat(release): verify Windows automatic updates end to end #3240/feat(win): Abort-path installer rollback with backup retention #3265 heads) are plausibly this same product-level startup failure — the harness could not see app stderr until now, so they were attributed to CDP attach. The capture is in place going forward, so future occurrences will self-identify.
    2. The 45s election deadline plus a wedged named-pipe endpoint on a slow runner is enough to kill startup outright (the app exits rather than retrying the spawn). Whether the right fix is a longer deadline, a respawn-on-unresponsive retry, or endpoint health hygiene is a product decision — but the failure is user-visible (app fails to launch), not just CI noise.

    Evidence grade: run log + source read (connect-or-spawn.ts:37,340, startup-error.ts:52); not locally reproduced.

  3. liugddx commented on Aug 21, 2026

    @liugddx
    Member

    Status after #3382 and the latest #3265 Windows run:

    Recommended next product investigation:

    1. Run a repeated fresh-profile startup stress loop under CPU/disk pressure while capturing the execution candidate/Host child's stderr, exit status, endpoint publication time, first successful pipe response, and election deadline.
    2. Classify the failing cut before changing timeouts: child exited, child alive with a published-but-unresponsive pipe, or stale endpoint/registration observed by election.
    3. If retry is added, bind it to the exact manager-owned candidate/Host identity and prove cleanup before respawn; do not turn the 45s deadline into an unbounded wait or hide the fault with verifier retries.

    The verifier already preserves app-side fatal output, so future occurrences should stay attributable to this issue.

    中文:最新 #3265 Windows 全链路多次启动均通过,本轮未复现 host_unresponsive,但单次未复现不能关闭问题;全新 profile 已证明旧状态不是必要条件。下一步应做带 Host 子进程 stderr/退出码/pipe 时间线的压力复现,先定位 child exit、活着但 endpoint 无响应、还是 election 观察到 stale endpoint,再决定是否做有身份约束的单次 respawn。

  4. liugddx commented on Aug 21, 2026

    @liugddx
    Member

    @Joob1n I am starting a narrow product-side diagnostic slice for #3279 from current main, avoiding the verifier surface you previously claimed.

    Scope:

    • preserve candidate/Host child stderr + exit/endpoint/election timing as bounded structured startup evidence;
    • add a repeated fresh-profile startup stress entry point for Windows;
    • do not increase the election deadline or add retry/respawn until the failing cut is classified;
    • no changes to the release verifier's existing CDP/process-probe ownership unless a product diagnostic needs a minimal read seam.

    I will post the exact branch/PR and evidence before proposing a behavior fix. Please flag any active overlap and I will rebase the ownership boundary rather than duplicate work.

  5. Joob1n commented on Aug 21, 2026

    @Joob1n
    MemberAuthor

    @liugddx No overlap — go ahead, and thank you for asking rather than assuming.

    The verifier half I claimed on 2026-08-20 was capturing the Host's own stderr into the lifecycle failure output. That is already done, and you did it: #3327 landed it, and your comment above is reading the evidence it produces ("with the newly added stderr capture from #3327"). There is nothing left of my claim to protect.

    The product half I said I would attempt after the release-verifier PRs landed. Those have landed and I did not start it, so it is yours with nothing to rebase around. I have unassigned myself so the issue does not read as covered when it is not.

    Your scope looks right to me, and one part of it is the part I would have gotten wrong: not increasing the election deadline or adding retry/respawn until the failing cut is classified. A timeout that is widened before the cause is known stops reproducing without being fixed, and the evidence goes with it.

    One observation, offered as context and not as evidence — it is from macOS and this issue is Windows. While measuring something unrelated on my own machine, codesign --timestamp round trips varied between 285 ms and 2019 ms on the same binary, same machine, minutes apart. Nothing about that is specific to codesign; it is what an unremarkable network round trip looks like when the runner is not the constant it appears to be. If the fresh-profile stress loop can vary CPU/disk pressure between iterations rather than holding it steady, a startup deadline that only fails under a particular ratio of contention would surface faster than one measured at a fixed load. Discard it if the failing cut turns out to be deterministic.

    I will stay out of the way. If a product diagnostic ends up needing a read seam into the verifier's CDP or process-probe surface, ping me — I know that code and can review the seam quickly rather than have you work around it.

  6. Joob1n commented on Aug 21, 2026

    @Joob1n
    MemberAuthor

    Correction to my last comment: I said I had unassigned myself, and I could not. Both paths are refused for an outside contributor — ReplaceActorsForAssignable on the GraphQL side, and the REST assignees endpoint the same way.

    So the assignment still reads Joob1n and that is now wrong: nothing on this issue is mine, and the half I did claim was delivered by #3327. @liugddx, if you have the permission, please take the assignment — or any committer can drop me from it. Until then the field says this issue is covered when it is not, which is the opposite of what I want it to say.

  7. liugddx commented on Aug 21, 2026

    @liugddx
    Member

    Product-side diagnostic PR: #3454 (b2315f21eec3076592d86643d492c6e7f5e70183).

    Scope correction to my earlier kickoff comment: I wrote "child stderr + exit/endpoint/election timing," which was stronger than the narrow implementation and should not be read as the PR contract. The detached Candidate still uses stdio: ignore; #3454 does not take ownership of child stderr or claim exact endpoint publication/phase timestamps. Adding pipes here would change detached-process I/O and lifecycle before the failing cut is known.

    The actual bounded evidence is:

    • exact Candidate PID and startup-attempt identity;
    • running/exited/unknown state with exit code/signal when observed;
    • final election elapsed time and true process-attempt count (including coalesced launches);
    • endpoint-connected observation plus fixed election counters;
    • the last safe registration state/PID/lifecycle/generation summary.

    No root path or endpoint value is added to the startup error. The PR also adds a fresh-root Windows stress runner with per-iteration NDJSON and exact Candidate settlement. It does not change the 45s deadline, retry/respawn, backoff, endpoint behavior, or failure classification.

    On the exact PR head, a fresh-root Windows run connected 1/1 in 5.642s; the exact Candidate exited cleanly and the new root was removed. The PR is draft pending core CI and Release Windows check on the merged result.

  8. liugddx commented on Aug 21, 2026

    @liugddx
    Member

    CI follow-up for #3454 exact head b2315f21eec3076592d86643d492c6e7f5e70183:

    The PR is now ready for review. It remains a diagnostic slice, not a claim that #3279's startup behavior is fixed.

    I also attempted the assignment correction requested above, but GitHub rejected it with ReplaceActorsForAssignable; the issue still lists @Joob1n despite the explicit ownership handoff. A committer with assignment permission will need to remove or replace that stale field.

  9. liugddx commented on Aug 22, 2026

    @liugddx
    Member

    #3454 produced a classified packaged recurrence on exact head 8ab1ce2cd78f4f5fe7eadc5bd981b1aa4f65e00b.

    Release Windows check 32567384310, final "Verify automatic update end to end" startup:

    {
      "deadlineMs": 45000,
      "elapsedMs": 45174,
      "candidateLaunches": 69,
      "sawEndpointConnected": true,
      "observations": {
        "notRegistered": 73,
        "connectFailed": 0,
        "handshakeFailed": 0,
        "connected": 1,
        "readyWaitFailed": 1,
        "deadlineElapsed": 0
      },
      "lastRegistration": {
        "pid": 8268,
        "state": "recovering",
        "lifecycleMode": "ephemeral",
        "generation": "0.1.11"
      },
      "latestCandidate": {
        "pid": 9812,
        "startupAttemptId": "8edc95c0-e691-46c5-93ca-eefe5b837cc6",
        "state": "running"
      }
    }

    Classification:

    • not a pre-existing stale endpoint: this was a fresh automatic-update fixture and the endpoint was reached;
    • not a Candidate that exited before publication: the latest exact attempt was still running;
    • one Host accepted a connection but did not become ready before the election deadline; its last registration was recovering;
    • before registration stabilized, the same election made 69 distinct process attempts while observing not_registered 73 times.

    This gives a concrete behavior-fix path: keep at most one exact live Candidate attempt per election until that attempt exits or a registration/Host appears, then decide whether another launch is warranted. The fix must retain the existing bounded deadline and exact-attempt cleanup; it should not be implemented as a verifier retry or a wider timeout.

    I am not adding that behavior change to #3454 because the PR is intentionally diagnostic-only. Core CI and audit on the exact head passed; the Windows lane passed packaging, packaged smoke, and pinned upgrade/uninstall before this product-level recurrence.

  10. liugddx commented on Aug 22, 2026

    @liugddx
    Member

    Behavior fix PR: #3512 (814d95f663c74fb6bd804de5eed3ae37d2eae482).

    It addresses the concrete convergence failure captured by #3454 without expanding the diagnostic PR:

    • one election keeps at most one Candidate process in flight;
    • the gate opens only after that process's exact exited promise settles;
    • spawn rejection remains retryable because no process was created;
    • an unknown/rejected exit signal fails closed for that election;
    • the production 45s deadline, readiness polling, failure classification, endpoint behavior, and cleanup ownership are unchanged;
    • cross-client/global coalescing and full LaunchLease architecture remain out of scope.

    Direct evidence includes pending/exit/spawn-reject table paths, a real owned Candidate exit followed by exactly one production successor, and a real Electron Client reporting exactly one Candidate PID. The two startup source files now trigger Release Windows check, so packaged automatic-update evidence is required before the draft is opened for review.

    #3454 stays frozen until #3512 reaches a final disposition. If #3512 lands, #3454 can rebase, remove its Node-only stress harness, and close its separate counter-completeness P3 without mixing the behavior repair into the diagnostic PR.

  11. liugddx commented on Aug 22, 2026

    @liugddx
    Member

    #3512 exact-head CI is green:

    The PR is now ready for human review. This is one successful exact-head packaged run, so it verifies the repaired path but does not justify broader global/cross-client coalescing claims. #3454 remains frozen until #3512 is reviewed and merged; after that it can rebase and be reduced to the permanent diagnostic surface.

  12. liugddx commented on Aug 23, 2026

    @liugddx
    Member

    Final status after #3512 and the trimmed #3454 rerun:

    The issue now has a behavior fix merged and a permanent diagnostic PR ready for normal human review. It should remain open for any future recurrence, not be closed solely because one packaged run passed.

  13. added theissue type on Aug 29, 2026
  14. Astro-Han commented on Sep 17, 2026

    @Astro-Han
    Contributor

    The behavior fix landed in #3512, which keeps at most one Candidate in flight per election and opens the gate only after that process's exit settles; that is the direct cause of the 69-launch storm captured here. #3454 landed the permanent startup diagnostics, and there has been no reported recurrence since.

    Closing as fixed. Please reopen with a new classified capture if it happens again.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions