Skip to content

One component's hung install cleanup deadlocks startup indefinitely — no listener opens, nothing is logged, no timeout can release it (root cause: #2076) #2072

Description

@heskew

Revised (third mechanism revision). Earlier revisions of this issue attributed the hang to (1) the 1-hour per-spawn timeout, (2) awaiting 'close' without 'exit' with a grandchild holding inherited pipes, and (3) a self-owned component-preparation lock whose wait deadline is renewed forever. All three are refuted — see "Retracted explanations" below and the comment history for the trail. The mechanism below is verified against both the running bundled core and runtime inspector state.

Summary

A component that enters preparation on boot and whose install spawn leaves an unreaped child makes startup deadlock indefinitely. The node is completely unavailable throughout — harper cluster_status reports Harper is not running, no listener opens, replication peers cannot connect — and nothing is logged after the first second of boot. There is no timeout that can release it.

The underlying defect is an unbounded process-group termination poll, filed separately as #2076. This issue covers what makes it fatal rather than merely untidy: startup is gated on component preparation, so one component can take the whole node down, silently and permanently, on any restart.

Verified on @harperfast/harper-pro@5.2.0 as running in production; the same code shape is on main.

Observed

A config declared three package-based components. One had no lock entry in harper-application-lock.json and no installed directory, so it entered preparation on boot. Its package: was an SSH git remote.

The reason that install did not complete is undetermined, and is not required for this report. SSH deploy-key auth was configured: <rootPath>/ssh/ contained a config defining exactly the host alias used by that URL, with an IdentityFile pointing at a present 0600 OpenSSH private key, plus a populated known_hosts. That is the directory materializeGitSSH() consumes (join(getConfigValue(CONFIG_PARAMS.ROOTPATH), 'ssh')), and a legacy plaintext key of that form is explicitly supported. git 2.47.3 is installed. So this was not a missing-credentials case. (A hypothesis that ROOTPATH might be unset during a boot-time spawn was checked and judged unreachable — installApplications() initialises config before any spawn.)

After the last log line —

(node:1) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true ...

— nothing more was ever emitted. State at +90 minutes:

  • harper cluster_status → Harper is not running.
  • grep -c 'successfully started' → 0; grep -c 'timed out after' → 0
  • 19 threads, all state S (sleeping); CPU 0.76%, memory flat
  • the direct spawn child (sh) present as a zombie, PPid: 1, unreaped
  • one component-preparation lock ticket, published at boot, never released

From an inspector session on the wedged process (read-only; main thread, pid 1, uptime 5431 s):

  • threads global is an empty array — no worker threads started; server has no listener. The block is inside loadRootComponents → installApplications().
  • process._getActiveRequests() → empty.
  • libuv census: {async:9, timer:2, check:2, idle:1, prepare:1, pipe:2, signal:3, fs_event:2, loop:1} — no process handle, so the ChildProcess handle had already closed.
  • The two pipes are the process's own fd:1/fd:2, not a child's.
  • Exactly one persistent timer, same address across samples, re-arming at ~16-50 ms. No long timer of any kind.

Mechanism

The install spawn's 'close' handler awaits terminateProcessTree, which on the SIGKILL escalation path calls:

await waitForConfirmedTermination(() => processGroupIsAlive(processGroupId));

and in components/Application.ts:

const PROCESS_TERMINATION_POLL_MS = 25;

function processGroupIsAlive(processGroupId: number): boolean {
    try { process.kill(-processGroupId, 0); return true; }
    catch (error: any) { return error.code === 'EPERM'; }
}

export async function waitForConfirmedTermination(isAlive, pollMs = PROCESS_TERMINATION_POLL_MS) {
    while (await isAlive()) await delay(pollMs);   // no deadline, no cap
}

A process that has exited but not been reaped still occupies a PID, so kill(-pgid, 0) succeeds and processGroupIsAlive returns true for a group whose only remaining member is a zombie. SIGTERM/SIGKILL are no-ops against an already-dead process, and waitForConfirmedTermination has no deadline — unlike its sibling waitForProcessGroupExit(processGroupId, timeoutMs), which is bounded. So the await never completes.

This accounts for every observation: a ~25 ms re-arming poll and no long timer; the zombie child being precisely what the loop waits on; no process handle, because 'close' already fired and closed it; the lock ticket unreleased because the process is still inside the critical section that holds it; sleeping threads at ~0.76% CPU; and no output, indefinitely.

Full analysis of the primitive is in #2076.

Why the install path is entered at all

components/Application.ts skips installation only when all three hold:

existsSync(application.dirPath) &&
harperApplicationLock.applications[name] &&
JSON.stringify(harperApplicationLock.applications[name]) === JSON.stringify(applicationConfig)

A component that has never successfully installed satisfies none of them, so preparation re-runs on every boot. There is no "this failed before, don't gate startup on it again" state. (An empty directory satisfies existsSync, so a half-finished install reads as installed — a related trap.)

Why it takes the whole node down

  • server/loadRootComponents.js — if (isMainThread && !process.env.HARPER_SAFE_MODE) await installApplications();
  • server/threads/threadServer.js — loadRootComponents(true) is awaited; listenOnPorts() is only reached after it resolves.

No listener opens until every component's preparation resolves, so one component's hung cleanup takes the entire node out rather than degrading that component.

Peers cannot distinguish this from a network fault: they report Client network socket disconnected before secure TLS connect and loop their subscription reconciler indefinitely, because the replication port never opens.

It is also latent — the config entry sits inert until the next restart, so the outage surfaces long after the change that caused it and gets attributed to whatever triggered the restart.

Retracted explanations

Recorded so nobody re-treads them:

  1. "Blocks for the full 1-hour spawn timeout." Refuted — no long timer is ever armed; the node passed 90 minutes with nothing pending that could release it.
  2. "'close' awaited without 'exit', with a grandchild holding inherited pipes." Refuted — the libuv census shows only the process's own fd:1/fd:2 pipes and no process handle, so no child pipes were held open. 'close' did fire; the hang is inside its handler.
  3. "Waiting on its own component-preparation lock ticket, deadline renewed forever." Refuted — scanLiveClaims is called with owner.token and filters claim.owner?.token === ownToken, so a process cannot block on its own claim; blocker would be falsy and the wait loop would break immediately. The unreleased ticket is a symptom of hanging inside the critical section, not its cause. (Credit to an outside review for catching this.)

Expected

  1. Fix the unbounded poll — see waitForConfirmedTermination polls without a deadline, and processGroupIsAlive counts an unreaped zombie as alive — a spawn's cleanup can hang forever #2076 (bound the wait; do not count an unreaped-dead group member as alive).
  2. Do not gate the whole node on one component's preparation. Start without it and report the component as failed, as a failed component load already is.
  3. Bound component preparation independently, so no single primitive inside it can hold boot open indefinitely.
  4. Log progress. An operator currently sees an empty log and a dead node. "Preparing component X" / "still waiting on X after Ns" would make this diagnosable in seconds.
  5. Consider recording failed installs so a known-unresolvable package does not re-block every subsequent boot.

Workaround

HARPER_SAFE_MODE skips installApplications() entirely, so a node already wedged this way can be booted with it set. Otherwise the offending component entry must be removed from the config before restarting.

Reproducer

  1. Add a component entry whose package: points at a git remote whose clone will fail, with no install: block.
  2. Ensure it has no lock entry and no installed directory.
  3. Restart, and arrange for the spawn's direct child to exit without being reaped.

Expect: no log output after the shell-spawn warning, Harper is not running indefinitely, a ~25 ms poll as the only timer, and a zombie child.

Related

Filed from a field incident; cluster, host, component and repository identifiers omitted.

Activity

  1. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Confirmed the load-bearing claim — startup does block on component installation, verified at v5.2.0:

    • server/loadRootComponents.js:18 — if (isMainThread && !process.env.HARPER_SAFE_MODE) await installApplications();
    • server/threads/threadServer.js:180-181 — require('../loadRootComponents.js').loadRootComponents(true) is awaited, while listenOnPorts() is not reached until line 225.

    So no listener opens until every component's install resolves. That is what turns one unresolvable package: into a full-node outage rather than one degraded component.

    Note the !process.env.HARPER_SAFE_MODE guard on that line: setting HARPER_SAFE_MODE skips installApplications() entirely, which is a usable operator workaround for a node already wedged this way — and is also further evidence that the install step is what's gating boot.

  2. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Corrected line references. My originals were read from a working tree at 5.1.3; verified against origin/main (24f67d8). The substance is unchanged — only the locations move.

    Claim Correct location on main
    1-hour default components/Application.ts:481 — const DEFAULT_COMMAND_TIMEOUT_MS = 60 * 60 * 1000;
    spawn signature default components/Application.ts:1690-1695 — nonInteractiveSpawn(... timeoutMs: number = DEFAULT_COMMAND_TIMEOUT_MS ...)
    install path passes install?.timeout components/Application.ts:902, 964, 1022; also 1238 — application.install?.timeout ?? DEFAULT_COMMAND_TIMEOUT_MS
    shell: true, stdio: ['ignore','pipe','pipe'] components/Application.ts:1769
    resolves on 'close', rejects on 'error', no 'exit' listener components/Application.ts:1843 (error), 1860 (close)
    install-skip condition components/Application.ts:1389-1391

    So on current main the default is still one hour, now via a named constant, and there is still no 'exit' listener — a child that has exited cannot unblock the await if its stdio pipes are held open by a grandchild.

    The boot-blocking chain in my earlier comment (server/loadRootComponents.js:18, server/threads/threadServer.js:180-181 vs listenOnPorts() at 225) was verified at v5.2.0 and is unchanged.

  3. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    On terminateProcessTree — it does not mitigate this, and there is field evidence.

    An outside review raised the fair question of whether the existing process-group handling already covers this, i.e. whether awaiting 'close' alone is actually a defect. Checking it:

    • components/Application.ts:1775 — detached: process.platform !== 'win32'
    • terminateProcessTree(childProcess, closePromise) is invoked in two places: the timeout path (~1794) and inside the 'close' handler (~1871).

    Both reachable paths are downstream of either 'close' firing or the 1-hour timeout elapsing. terminateProcessTree is therefore a cleanup mechanism, not a settlement mechanism — it cannot cause the awaited promise to settle when 'close' never fires. The startup block is unaffected by it.

    Version boundary: terminateProcessTree is absent at v5.1.26 and present in v5.2.0 and main (8 occurrences each).

    The decisive evidence: the node this was observed on was running 5.2.0 — i.e. a build that already has terminateProcessTree — and startup was still blocked with zero log output, all threads sleeping, CPU ~0.7%, and Harper is not running for over an hour, with the direct sh child sitting as an unreaped zombie the whole time. So the existing group handling demonstrably does not prevent this.

    That keeps the requested fix as filed: settlement must not depend solely on 'close', and component preparation needs its own short bound independent of the 1-hour per-spawn default.

  4. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Correcting myself: the falsifiable prediction in the body is REFUTED, and the real bound is worse than filed.

    I predicted the spawn's 1-hour timeout would fire at spawn + 3,600,000 ms and startup would then proceed or fail. It did not. At +79 minutes the affected node showed:

    • grep -c 'timed out after' → 0
    • grep -c 'successfully started' → 0
    • log still frozen at the DEP0190 shell-spawn warning
    • harper cluster_status → still Harper is not running

    So whatever is blocking is not settling on the per-spawn 1-hour timer, and citing "one hour" as the outage length was wrong.

    The governing bound for this path is the component-preparation lock wait, not the spawn timeout. components/Application.ts:1239-1290 wraps preparation in withComponentPreparationLock(...) with:

    timeoutMs: MAX_GIT_EXTRACTION_COMMANDS * DEFAULT_COMMAND_TIMEOUT_MS
             + MAX_INSTALL_COMMANDS * commandTimeoutMs
             + COMPONENT_PREPARATION_WAIT_MARGIN_MS
    

    With DEFAULT_COMMAND_TIMEOUT_MS = 3600000 (:481), COMPONENT_PREPARATION_WAIT_MARGIN_MS = 30000 (:482), MAX_GIT_EXTRACTION_COMMANDS = 4 (:483), MAX_INSTALL_COMMANDS = 2 (:484), that is 21,630,000 ms ≈ 6 hours for a git-sourced component. Please read the impact section as up to ~6 hours of total node unavailability per restart rather than one hour.

    What is now unresolved: why the spawn's own 1-hour timer did not fire, given DEP0190 proves spawn() was reached and the timer is armed immediately afterward. Two candidates, neither confirmed:

    1. That spawn settled and something later is blocking — but the direct sh child remains an unreaped zombie, which argues against clean settlement.
    2. Preparation is waiting to acquire a lock whose recorded owner is this same process/thread. The observed lock ticket was {"pid":1,"threadId":0,...} and isOwnerAlive (:1300) is owner.pid !== process.pid || isThreadRunning(owner.threadId) — for a ticket owned by the current pid it reduces to isThreadRunning(0), so a live main thread reads as a live holder. If the waiter is that same thread, it waits on itself until the ~6h budget expires.

    Distinguishing these needs an inspector session on the wedged process, which I did not have access to. I would rather leave the mechanism explicitly open than assert the wrong one twice.

    Unchanged and still verified: startup is gated on component installation (server/loadRootComponents.js:18, server/threads/threadServer.js:180-181 before listenOnPorts() at 225); a component with no lock entry and no installed directory re-enters preparation on every boot (:1389-1391); the node is fully unavailable and logs nothing for the entire window. The headline defect — an unresolvable component package silently taking a whole node out of service on restart — stands; only my timing claim was wrong.

  5. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Mechanism resolved via inspector session on the wedged process. It is unbounded — neither of my earlier timing claims was right, and the real cause is a self-owned preparation lock whose wait deadline renews forever.

    Runtime evidence (read-only introspection, main thread, pid 1, uptime 5431s)

    • threads global is an empty array — no worker threads started; server has no listener. Confirms the block is inside loadRootComponents → installApplications().
    • process._getActiveRequests() → empty. No pending I/O request.
    • libuv handle census: {async:9, timer:2, check:2, idle:1, prepare:1, pipe:2, signal:3, fs_event:2, loop:1} — no process handle, so no live child process.
    • The two pipes are the process's own fd:1/fd:2 (stdio: "stdout"/"stderr"), not a child's. The "grandchild holds the child's pipes" theory is therefore wrong.
    • Exactly one persistent timer, same address across samples, re-arming at ~16-50 ms. No long timer exists — no 1-hour spawn timer, no multi-hour lock timer.

    Why no timeout will ever fire

    components/componentPreparationLock.ts:

    • LOCK_POLL_INTERVAL_MS = 50 (:13) and await delay(LOCK_POLL_INTERVAL_MS) (:295) — matches the observed ~50 ms timer exactly.
    • The wait bound is a performance.now() deadline (:251), not a setTimeout — which is why no long timer is armed.
    • :283-293 — on deadline expiry, if the holder's liveness is confirmed the deadline is renewed:
    if (performance.now() >= deadline) {
        if (await ownerLivenessConfirmed(blocker, options)) {
            deadline = performance.now() + timeoutMs;   // renewed
        } else { throw new Error(`Timed out waiting for component preparation lock ...`); }
    }
    • ownerLivenessConfirmed delegates to options.isOwnerAlive, which at components/Application.ts:1300 is:
    isOwnerAlive: (owner) => owner.pid !== process.pid || isThreadRunning(owner.threadId)

    The observed lock ticket was {"pid":1,"threadId":0,...}, in a process where process.pid === 1. So owner.pid !== process.pid is false and it reduces to isThreadRunning(0) — the main thread, which is running, because it is the thread doing the waiting. Liveness is "confirmed", the deadline is renewed, and the loop polls every 50 ms forever.

    A thread cannot be both the holder and the waiter, but isOwnerAlive cannot distinguish those cases, so a stale ticket recorded against the current pid/thread produces a permanent, silent, self-inflicted deadlock. The existing comment at :89-92 shows the renew-forever hazard was anticipated; the stricter check still delegates to the same predicate, so it does not catch this case.

    Also waitingReported (:279-281) fires onWait once, so there is no repeating log to reveal the wait.

    Corrections to this issue

    • "blocks for the full 1-hour spawn timeout" — wrong; that timer is never armed.
    • My later "~6 hours via the preparation-lock budget" — also wrong; the budget is renewed rather than enforced.
    • Correct statement: startup blocks indefinitely. The observed node passed 90 minutes with no timer that could ever release it, and would not have recovered without intervention.

    Suggested fix

    isOwnerAlive should return false when the recorded owner is the current process and the current thread — that combination is impossible for a genuine live holder and means the ticket is stale. Independently, deadline renewal should be bounded (renew at most N times, or cap total wait) so no liveness predicate can produce an unbounded wait.

    Retitling suggestion: the "1-hour spawn timeout" framing in the title is inaccurate; this is an indefinite startup deadlock on a self-owned component-preparation lock.

  6. changed the title [-]A component package that cannot install blocks startup for the full 1-hour spawn timeout with no log output; node reports 'Harper is not running' throughout[/-] [+]Startup deadlocks indefinitely waiting on a self-owned component-preparation lock (wait deadline renewed forever); one unresolvable component package takes the whole node down[/+] on Aug 4, 2026
  7. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Exposure window — this code is days old, which is likely why it has not been hit before.

    components/componentPreparationLock.ts is absent from every OSS release tag checked — v5.1.19, v5.1.23, v5.1.24, v5.1.26, and v5.2.0. It exists only on main, introduced by 5c6b68471 "Fix concurrent component installation" (the concurrent-install/node_modules corruption work also referenced from #1996).

    It nevertheless is running in the field: verified inside the affected container that @harperfast/harper-pro@5.2.0 bundles a core snapshot containing the file, dated Aug 1 02:21, with the relevant lines at the same offsets as main (LOCK_POLL_INTERVAL_MS = 50 :13, deadline :251, renewal :284-287, poll :295). So the pro build was cut from a core commit later than the OSS v5.2.0 tag. Anyone reproducing from the OSS tag will not find this file — check the bundled core in the running package instead.

    Why the trigger is narrow, which compounds the low exposure:

    1. Only package-based components are considered at all — Application.ts:588 skips config entries without 'package' in applicationConfig. Payload-deployed components (the usual deploy_component path) never enter installApplications.
    2. A component that has installed successfully once has both a directory and a matching lock entry, so :1389-1391 skips it on every subsequent boot — it never touches the preparation lock again.
    3. The vulnerable window is therefore narrow: a package-based component that has never successfully installed, on a boot. With a working package URL that resolves on first boot and is skipped thereafter.
    4. It additionally needs the package to be permanently unresolvable (here: an SSH remote with no credentials in the container), and restarts are infrequent — the bad entry sat inert until an upgrade forced one.

    Net: affected builds are recent pro releases carrying a post-v5.2.0 core snapshot, and the trigger requires a never-installed package-based component plus a restart. That is consistent with this being an early sighting rather than a long-standing silent problem — but it also means exposure grows as that core snapshot propagates into more releases.

  8. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Correction to the "no credentials" claim in the Observed section (body updated). My original wording said the container had no credentials, evidenced by ~/.ssh being absent — that check was run via docker exec as root, so ~ expanded to /root/.ssh, which is not where Harper's runtime user lives. Wrong path, and the blanket claim was too strong.

    Re-checked properly. The corrected position:

    • No .ssh exists for either the container root or the harperdb user, so there is genuinely no SSH key, known_hosts, or host-alias config.
    • There is a git credential mechanism — gitCredentialHelper.js / gitCredentialServer.ts, via HARPER_GIT_CREDENTIAL_SOCKET, wired as credential.helper or GIT_ASKPASS. But its own comments show it answers HTTPS prompts (Username for 'https://github.com'), i.e. token/password auth — not SSH key auth, so it cannot serve an SSH remote.
    • getGitSSHCommand (which sets GIT_SSH_COMMAND on main) is not in this build.
    • Credentials are supplied per deploy request; a config-declared package: installed at boot has no request to carry them.
    • git itself is present (2.47.3), so the failure is authentication, not a missing binary.

    So the conclusion stands — that URL could not have authenticated — but for a more specific reason than originally stated: the instance has an HTTPS credential channel and the component was configured with an SSH remote.

    One check I am explicitly withdrawing: I attempted to read /proc/1/environ to enumerate live SSH_*/GIT_* variables and it failed with permission denied. Any inference that no such variables are set is unsupported — I could not read the process environment.

    None of this affects the deadlock mechanism, which is in lock acquisition and independent of why the install could not succeed.

  9. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Retracting the credentials claim entirely — it was wrong. SSH deploy-key auth was configured on the affected node.

    I twice asserted the instance had no way to authenticate the SSH remote. Both checks looked at ~/.ssh, which is not a path Harper uses. The mechanism is materializeGitSSH() (components/Application.ts ~1326-1345), which reads join(getConfigValue(CONFIG_PARAMS.ROOTPATH), 'ssh') — a durable <rootPath>/ssh directory whose *.key files are materialized into a transient 0700 dir per git-over-SSH spawn, with the ssh config copied and its IdentityFile lines repointed at the transient copies.

    On the affected node that directory contained:

    • a config defining exactly the host alias used in the failing package URL, with IdentityFile pointing at
    • a present 0600 OpenSSH private key (plaintext, i.e. a legacy unsealed key — explicitly supported per the function's own docs: "Legacy plaintext keys … are copied through unchanged"), and
    • a populated known_hosts.

    So the trigger was not missing credentials. I have removed that claim from the body. The reason the install did not complete is now undetermined and, importantly, is not load-bearing for this issue — the deadlock is in lock acquisition and is independent of why the install failed.

    One candidate a maintainer may want to look at, explicitly unproven: materializeGitSSH() returns undefined when ROOTPATH is unset, its comment noting "config not initialized (e.g. an install-time spawn) — no ssh dir to read". If a boot-time installApplications() spawn runs before that config value is available, no GIT_SSH_COMMAND is exported and git falls back to a ~/.ssh that does not exist on this image — an auth failure despite correctly configured keys. That would be a distinct bug.

    Also correcting two smaller errors from my earlier comment: getGitSSHCommand is not a symbol in either the running build or main (it came from an out-of-date local tree), and both trees do contain GIT_SSH_COMMAND handling — so my statement that this build has no SSH-command path was false. The HTTPS credential-helper machinery (gitCredentialHelper.js / gitCredentialServer.ts) is real but simply orthogonal to SSH remotes.

    Apologies for the churn on the trigger description. The mechanism section — verified against the running bundled core and against runtime inspector state — is unaffected.

  10. changed the title [-]Startup deadlocks indefinitely waiting on a self-owned component-preparation lock (wait deadline renewed forever); one unresolvable component package takes the whole node down[/-] [+]One component's hung install cleanup deadlocks startup indefinitely — no listener opens, nothing is logged, no timeout can release it (root cause: #2076)[/+] on Aug 4, 2026
  11. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Mechanism corrected (third revision) and root cause split out to #2076.

    An outside review refuted the self-owned-lock explanation, and it was right. scanLiveClaims is invoked with owner.token from inside the wait loop and filters claim.owner?.token === ownToken, so a process cannot block on its own claim — blocker would be falsy and the loop would break immediately. The unreleased ticket is therefore a symptom of hanging inside the critical section, not the cause. That retraction is now in the body.

    Following the evidence instead of the hypothesis produced a mechanism that accounts for every observation: the 'close' handler awaits terminateProcessTree, whose SIGKILL escalation path calls waitForConfirmedTermination(() => processGroupIsAlive(pgid)). processGroupIsAlive probes with process.kill(-pgid, 0), which succeeds for an exited-but-unreaped process, and waitForConfirmedTermination has no deadline (unlike the bounded sibling waitForProcessGroupExit(pgid, timeoutMs)). With a zombie child in the group, the poll runs every PROCESS_TERMINATION_POLL_MS = 25 forever.

    That matches the measured state exactly — one ~16-50 ms re-arming timer and no long timer, no process handle (so 'close' had fired), a zombie child as the thing being waited on, the lock ticket held because we are inside the critical section, sleeping threads at 0.76% CPU, and no output.

    The primitive is filed on its own as #2076, since it can hang any caller of terminateProcessTree, not just component installs. This issue now covers what makes it fatal: startup is gated on component preparation, so one component takes the whole node down silently and permanently on any restart.

    Also settled: the ROOTPATH-unset hypothesis for the git failure is unreachable — installApplications() initialises config before any spawn — so the trigger for the failed install remains undetermined, and deliberately so; the deadlock does not depend on it.

    Three superseded explanations are listed explicitly in the body so they are not re-tried.

  12. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Both open questions from this issue are now resolved, while the node was still wedged.

    1. Why the child was never reaped — answered in #2076. Short version: detached: true makes the direct child a group leader; it spawned a grandchild; the direct child exited and was reaped normally (that is what fired 'close'); the grandchild was orphaned and re-parented to pid 1, which is node itself in a container with no tini/dumb-init. Node only waits on children it spawned via a libuv handle, and pid 1's SigCgt does not include SIGCHLD, so the orphan is never reaped. Its group's only remaining member is that zombie, which is exactly what keeps kill(-pgid, 0) succeeding forever. So the precondition for #2076 is structural for containerized deployments, not a rare accident.

    2. Why the install "failed" — the premise appears to be wrong, and I am withdrawing it. There is no evidence the install failed:

    • The durable ssh config was complete and correct: HostName github.com, User git, IdentityFile <key>, IdentitiesOnly yes, with matching known_hosts entries. (An earlier idea that the host alias might not resolve is void — HostName is set, so the alias never needs DNS.)
    • materializeGitSSH() demonstrably ran: its transient 0700 dir was still on disk with the materialized 0600 key and rewritten config, mtime matching the spawn to the millisecond. So a working GIT_SSH_COMMAND was exported.
    • Whether the git command itself succeeded or failed is unknowable from outside the process: nonInteractiveSpawn buffers stdout/stderr and only prints them once the promise settles, which never happened.

    So the observable defect is the cleanup hang, not a clone failure. The component directory was never created, but that is equally consistent with a clone that completed into staging and then hung before being moved into place.

    Body updated accordingly: this issue no longer asserts the install failed, only that preparation was entered and never completed — which is all the report needs, since the deadlock is in spawn cleanup and is independent of the git outcome.

    One incidental finding worth noting for whoever fixes this: the transient decrypted-key directory outlives its documented lifetime while the spawn is hung (details in #2076). For nodes using sealed enc:v1: keys that turns an at-rest-encrypted key into an indefinitely-persisted plaintext file.

  13. Devin-Holland commented on Sep 24, 2026

    @Devin-Holland
    Member

    Second independent field occurrence, on 5.2.13 — reproduced three times on one node inside an hour, with two different package URLs. The mechanism section here holds exactly.

    Different cluster, different customer, ~7 weeks after this was filed. I hit this cold during an incident investigation and arrived at the same place from the operator side before finding this issue. Posting the independent confirmation plus four things that are new.

    Everything below is from the live node. Cluster, host, component and repository identifiers omitted, matching the convention in the body.

    Confirmation — identical signature on @harperfast/harper-pro@5.2.13

    Two nodes, node A and node B. Node A entered preparation on boot for a package-based component with no lock entry and no installed directory, and wedged:

    harper get_status            -> "Harper is not running."
    openssl s_client 127.0.0.1:9926  (inside the container) -> ECONNREFUSED
    openssl s_client 127.0.0.1:9933  (inside the container) -> ECONNREFUSED
    <rootPath>/operations-server -> does not exist
    

    Last line ever emitted, matching the body verbatim:

    (node:1) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true ...
    

    Process table, ~19 minutes in:

    UID   PID  PPID  C STIME TTY  TIME     CMD
    harper  1     0  0 13:52 ?    00:00:04 node /home/harperdb/.npm-global/bin/harper
    harper 281    1  0 13:53 ?    00:00:00 [git] <defunct>
    

    CPU 0.24%, RSS flat at 225 MiB, 28 threads with 1 running / 26 sleeping / 1 zombie. One unreleased preparation ticket, holder already dead:

    {"pid":1,"threadId":0,"processInstanceId":"d3a5dfdd-…","token":"ffc37f83-…","ticket":1}

    So: zombie in the group, no listener, no output, flat CPU, ticket held. The terminateProcessTree → waitForConfirmedTermination → processGroupIsAlive chain in the body accounts for all of it. One small difference from the original report — here the direct child left as a zombie is git itself rather than an intermediate sh, and it is parented to pid 1 (node, no init). Same structural precondition as described in the #2076 cross-post; it does not need a grandchild to reproduce.

    components/.deploy-aside/ was empty, so there was also no prior version to fall back to.

    Why this still reproduced: #2085 fixed the predicate on main only

    The obvious question is why this recurred at all when #2085 ("avoid startup hangs on zombie process groups") merged 2026-08-27. Answer: it is on main only and was never cherry-picked to v5.2.

    ref version isProcessGroupAlive in components/Application.ts
    main 5.3.0-beta.2 present (import + call)
    v5.2 5.2.13 absent

    Confirmed a second way, by reading the bundled core actually running on the wedged node (@harperfast/harper-pro@5.2.13, current v5.2 head). Both the TypeScript source and the compiled dist still carry the pre-fix body verbatim:

    // dist/core/components/Application.js:1649  — as shipped in 5.2.13
    function processGroupIsAlive(processGroupId) {
        try { process.kill(-processGroupId, 0); return true; }
        catch (error) { return error.code === 'EPERM'; }
    }

    isProcessGroupAlive does not exist anywhere in that build's manageThreads.js. I could not find a cherry-pick PR of #2085 against v5.2.

    So every 5.2.x node — including the newest, 5.2.13 — still counts an unreaped zombie as a live group member, and 5.3.0-beta.2 is the only release line carrying the fix.

    Would #2085 have prevented this?

    Reading the merged diff: yes, for this failure mode — though via the predicate rather than a timeout, which the PR description makes easy to miss. The description says waitForConfirmedTermination "is untouched and stays deliberately unbounded after SIGKILL", which is true of the loop, but the same PR rewired the predicate that loop polls:

     function processGroupIsAlive(processGroupId: number): boolean {
    -	try {
    -		process.kill(-processGroupId, 0);
    -		return true;
    -	} catch (error: any) {
    -		return error.code === 'EPERM';
    -	}
    +	return isProcessGroupAlive(processGroupId);
     }

    and on main isProcessGroupAlive falls through to a /proc scan when the leader is a zombie, requires two consecutive all-zombie snapshots, then returns false. For our group — whose only remaining member is a zombie git — the predicate flips to false, the unbounded loop exits, terminateProcessTree resolves, and boot proceeds. That is #2076's root cause fixed at the source, so the unbounded wait stops mattering for the zombie case.

    But it does not close this issue

    #2085 removes the trigger observed here; it does not remove the fatality. After it:

    So the failure class — one component silently and permanently taking a node down on restart — survives #2085 intact; only this particular path into it is closed, and only on 5.3. That seems worth keeping this issue open for on its own terms.

    Ask

    A cherry-pick of #2085 to v5.2 would stop this recurring on the 5.2 line, where the deployed fleet actually is. Happy to open it if that is wanted.

    New 1 — the git outcome really is irrelevant, demonstrated by substitution

    The body withdrew the claim that the install failed and called the trigger undetermined. This run supports that directly, because I got to watch two different package identifiers wedge the same way.

    The first attempt used a genuinely malformed identifier — a GitHub web tree URL (https://github.com/<org>/<repo>/tree/<branch>), which is not an installable spec at all — with no credentials, on a node that additionally logged:

    [main/0] [error]: SSH key <name>.key is encrypted but no secret custody is registered on this node; skipping it
    

    and carried secretCustody: {}. That is about as unambiguously doomed as a clone gets, and I initially wrote it up as the cause.

    The operator then redeployed with a correct spec — git+https://github.com/<org>/<repo>.git#<branch> — and the node wedged identically: same [git] <defunct>, same terminal DEP0190 line, same absent UDS, same flat CPU. Whatever the fix is, it cannot be "validate the package identifier": a well-formed one deadlocks the same way. Worth keeping in mind against any temptation to close this with input validation.

    New 2 — why peers cannot tell this from a network fault, mechanically

    The body notes peers report Client network socket disconnected before secure TLS connect and loop their reconciler. The reason is structural and worth recording, because it actively misdirects the on-call.

    On these hosts nginx listens on 0.0.0.0:9933 and proxies to the container's published replication port. nginx accepts the peer's TCP connection whether or not anything is behind it. So from node B:

    CONNECTED(00000003)
    no peer certificate available
    SSL handshake has read 0 bytes and written 358 bytes
    error:0A000126:SSL routines:ssl3_read_n:unexpected eof while reading
    

    — a TCP accept followed by a close before any TLS bytes come back. Meanwhile the same port probed from inside the wedged container is plain ECONNREFUSED.

    The peer therefore never sees a refusal, only a half-open-then-dropped connection, which reads as a TLS/CA problem or a flaky link. On this incident that cost real time: the symptom looks a lot like the peer-CA trust class of bug, and the surviving node's logs are full of Reconciling N wedged subscription(s) … lastCloseCode: 1006 climbing forever. Useful triage rule for whoever writes the runbook: probe the port from inside the target's own container. Refused there while the peer reports a pre-TLS drop means "no listener at all / boot never completed", not a TLS problem.

    New 3 — hdb_deployment is an already-wired progress surface that this leaves permanently stuck

    Coming in via deploy_component rather than a hand-edited config, there is a system-table row per attempt, and it shows the hang from the outside with no inspector needed:

    project:            <name>
    package_identifier: <url>
    status:             pending
    phase:              prepare
    event_log:          [ { t: <boot+Ns>, event: "phase", data: { phase: "prepare", status: "start" } } ]
    started_at:         <t>          completed_at: null      error: null
    

    Exactly one event. No second event, no error, no completion — 45+ minutes and counting, and the row replicates cluster-wide, so every node advertises a permanently pending deployment. After the redeploy there were two such rows stuck side by side.

    This seems directly relevant to Expected #4 and #5. The surface already exists and is already replicated; it just never gets a prepare heartbeat, a deadline, or a terminal state. A bounded prepare that writes status: failed here would satisfy "log progress" and "record failed installs so a known-unresolvable package does not re-block every subsequent boot" without inventing a new mechanism. It would also have made this diagnosable from the API in seconds.

    Related: credentials: None and restart_mode: rolling are both on the row, so the deploy path has the metadata to decide whether a given install can possibly succeed before it gates boot on it.

    New 4 — what separates a wedged node from a healthy one is the deploy method, not the node

    Correcting an earlier draft of this comment: I first attributed the asymmetry to node A being a fresh clone of node B. That is wrong, and the real answer isolates the trigger more usefully.

    The two nodes carry byte-identical SSH material — same <rootPath>/ssh/config, same sealed <rootPath>/ssh/*.key (matching sha256) — and both are correctly provisioned with the cluster's shared secret-custody keypair. Node A is not misconfigured.

    It nonetheless logs, during the boot-time install:

    [main/0] [error]: SSH key <name>.key is encrypted but no secret custody is registered on this node; skipping it
    

    This is an ordering bug, and it is reachable on any node. server/loadRootComponents.js:

    async function loadRootComponents(isWorkerThread = false) {
    	try {
    		if (isMainThread && !process.env.HARPER_SAFE_MODE) await installApplications();   // line 18
    	} catch (error) { ... }
    	// ...
    	await loadComponentDirectories(loadedComponents, resources);                          // line 36

    installApplications() is the first statement. Secret custody is itself a built-in component (secretCustody=@/dist/security/keyCustody.js in the built-ins list), and its startOnMainThread — which calls registerSecretCustody() — does not run until loadComponentDirectories() at line 36. So for the entire duration of a boot-time component install, getSecretDecryptor() returns undefined and every sealed enc:v1: key is skipped, no matter how correctly custody is provisioned. A runtime deploy, after components are loaded, has custody available; a boot-time install never does.

    The healthy node never logs this only because its components are already installed (lock entry + directory present), so no install spawn runs at boot.

    This is worth flagging because the same shape of hypothesis appears in the comment history and was cleared: materializeGitSSH() returning undefined when ROOTPATH is unset during an install-time spawn, ruled unreachable because installApplications() initialises config before any spawn. That reasoning is correct for ROOTPATH and does not carry over to custody, which is registered by a component rather than by config init, and therefore really is absent at that point.

    It is, however, not what wedged node A here — both failing deploys used HTTPS remotes, and materializeGitSSH() runs unconditionally before every git spawn regardless of protocol, so the skipped SSH key is logged but never consulted. It is a red herring that happens to log immediately before the spawn warning. Recording it because this issue's history shows how easily this path gets misattributed.

    What actually separates them is how each component arrived. Four deployment rows on this cluster:

    project package_identifier payload_size status
    app (working) none 189,869,451 success
    status-check git+https://…/status-check-fabric.git#semver:v1.0.0 (public) none success
    dev attempt 1 web tree URL (private) none pending / prepare
    dev attempt 2 git+https://…​.git#develop (private) none pending / prepare
    • Payload deploys never enter this path at all. The working application was uploaded as a ~181 MB payload with no package identifier — no git spawn, so no process group, so nothing to poll. Immune by construction.
    • A package deploy against a public repo also succeeds — status-check installs from a git URL on these same nodes with the same unusable key, because it needs no credentials.
    • Both wedges are package deploys against a private repo with credentials: None, and both originated on node A.

    So the precondition is narrow and clear: a package-identifier deploy whose git invocation cannot authenticate. Not clone state, not node identity, not a missing key. Node B has the same sealed-and-undecryptable key and the same absent custody; by inspection it would wedge the same way if the same deploy were issued against it (not tested — it is a customer node).

    One clone-related observation does survive: a newly cloned node has an empty .deploy-aside, so there is no previously-working copy to fall back to when preparation strands the component directory.

    • The body calls the failure latent because the entry sits inert until the next restart. Worth making explicit who writes that entry, and when: the deploy handler itself does, in components/operations.js — await configUtils.addConfig(req.project, applicationConfig) — keyed on the deploy request's project name rather than the package name. It is written as part of the deploy, before the component is known to be installable. So a deploy whose preparation never completes still leaves behind a config entry that arms the next boot. I removed the entry by hand and the node booted cleanly — listeners up, UDS present, get_status: Available. A redeploy of the same project four minutes later wrote it straight back, and the next restart wedged again. Manual removal is therefore durable only until the next deploy attempt of that project; every retry re-arms the wedge. That ordering — persist the config entry first, find out whether it can install second — looks like a cheap place to intervene alongside Expected Add a CODEOWNERS file to set PR reviewers #5.

    Timeline on one node, all within about 35 minutes: wedged on boot → exited → config entry removed by hand → booted clean and served → redeploy re-added the entry → restart → wedged again, [git] <defunct>, no UDS, DEP0190 as the last line.

    Endorsing Expected #2

    Of the five, item 2 is the one that would have contained this incident entirely. A component that cannot be fetched taking down replication, the operations API, the HTTP listener and the status endpoint — on a node whose other component was healthy and whose peer was fine — turns a bad deploy into a total outage with no surface left to report it. The node also silently stops answering the status check that gates GTM, so it drops out of rotation with no signal as to why.

    Happy to run anything else against this cluster while it is still in this state, or to re-run a specific probe if a maintainer wants a particular piece of inspector state captured.

  14. Devin-Holland commented on Sep 24, 2026

    @Devin-Holland
    Member

    The custody-ordering finding in my comment above is now filed on its own as #2780 — secret custody registers after installApplications(), so sealed enc:v1: SSH keys are always skipped during a boot-time component install, on any node regardless of provisioning.

    Filing it separately because it is a distinct defect with its own fix, not a manifestation of this one: it fails the install, and it is this issue's startup gating plus #2076's unbounded poll that turn that failed install into a dead node. It also reaches further than git — anything relying on a sealed secret during boot-time component preparation is affected.

    Noted there, and repeating here so it isn't re-tried: this is not the ROOTPATH-unset hypothesis that was correctly ruled out in this thread. That reasoning holds for ROOTPATH (config is initialised before any spawn) but does not carry over to custody, which is registered by a component rather than by config init.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    Priority

    P1

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions