Repository navigation
One component's hung install cleanup deadlocks startup indefinitely — no listener opens, nothing is logged, no timeout can release it (root cause: #2076) #2072
Description
Activity
Confirmed the load-bearing claim — startup does block on component installation, verified at v5.2.0:
server/loadRootComponents.js:18—if (isMainThread && !process.env.HARPER_SAFE_MODE) await installApplications();server/threads/threadServer.js:180-181—require('../loadRootComponents.js').loadRootComponents(true)is awaited, whilelistenOnPorts()is not reached until line 225.
So no listener opens until every component's install resolves. That is what turns one unresolvable
package:into a full-node outage rather than one degraded component.Note the
!process.env.HARPER_SAFE_MODEguard on that line: settingHARPER_SAFE_MODEskipsinstallApplications()entirely, which is a usable operator workaround for a node already wedged this way — and is also further evidence that the install step is what's gating boot.Corrected line references. My originals were read from a working tree at 5.1.3; verified against
origin/main(24f67d8). The substance is unchanged — only the locations move.Claim Correct location on main1-hour default components/Application.ts:481—const DEFAULT_COMMAND_TIMEOUT_MS = 60 * 60 * 1000;spawn signature default components/Application.ts:1690-1695—nonInteractiveSpawn(... timeoutMs: number = DEFAULT_COMMAND_TIMEOUT_MS ...)install path passes install?.timeoutcomponents/Application.ts:902,964,1022; also1238—application.install?.timeout ?? DEFAULT_COMMAND_TIMEOUT_MSshell: true,stdio: ['ignore','pipe','pipe']components/Application.ts:1769resolves on 'close', rejects on'error', no'exit'listenercomponents/Application.ts:1843(error),1860(close)install-skip condition components/Application.ts:1389-1391So on current
mainthe default is still one hour, now via a named constant, and there is still no'exit'listener — a child that has exited cannot unblock the await if its stdio pipes are held open by a grandchild.The boot-blocking chain in my earlier comment (
server/loadRootComponents.js:18,server/threads/threadServer.js:180-181vslistenOnPorts()at 225) was verified atv5.2.0and is unchanged.On
terminateProcessTree— it does not mitigate this, and there is field evidence.An outside review raised the fair question of whether the existing process-group handling already covers this, i.e. whether awaiting
'close'alone is actually a defect. Checking it:components/Application.ts:1775—detached: process.platform !== 'win32'terminateProcessTree(childProcess, closePromise)is invoked in two places: the timeout path (~1794) and inside the'close'handler (~1871).
Both reachable paths are downstream of either
'close'firing or the 1-hour timeout elapsing.terminateProcessTreeis therefore a cleanup mechanism, not a settlement mechanism — it cannot cause the awaited promise to settle when'close'never fires. The startup block is unaffected by it.Version boundary:
terminateProcessTreeis absent at v5.1.26 and present in v5.2.0 andmain(8 occurrences each).The decisive evidence: the node this was observed on was running 5.2.0 — i.e. a build that already has
terminateProcessTree— and startup was still blocked with zero log output, all threads sleeping, CPU ~0.7%, andHarper is not runningfor over an hour, with the directshchild sitting as an unreaped zombie the whole time. So the existing group handling demonstrably does not prevent this.That keeps the requested fix as filed: settlement must not depend solely on
'close', and component preparation needs its own short bound independent of the 1-hour per-spawn default.Correcting myself: the falsifiable prediction in the body is REFUTED, and the real bound is worse than filed.
I predicted the spawn's 1-hour timeout would fire at spawn + 3,600,000 ms and startup would then proceed or fail. It did not. At +79 minutes the affected node showed:
grep -c 'timed out after'→ 0grep -c 'successfully started'→ 0- log still frozen at the
DEP0190shell-spawn warning harper cluster_status→ stillHarper is not running
So whatever is blocking is not settling on the per-spawn 1-hour timer, and citing "one hour" as the outage length was wrong.
The governing bound for this path is the component-preparation lock wait, not the spawn timeout.
components/Application.ts:1239-1290wraps preparation inwithComponentPreparationLock(...)with:timeoutMs: MAX_GIT_EXTRACTION_COMMANDS * DEFAULT_COMMAND_TIMEOUT_MS + MAX_INSTALL_COMMANDS * commandTimeoutMs + COMPONENT_PREPARATION_WAIT_MARGIN_MSWith
DEFAULT_COMMAND_TIMEOUT_MS = 3600000(:481),COMPONENT_PREPARATION_WAIT_MARGIN_MS = 30000(:482),MAX_GIT_EXTRACTION_COMMANDS = 4(:483),MAX_INSTALL_COMMANDS = 2(:484), that is 21,630,000 ms ≈ 6 hours for a git-sourced component. Please read the impact section as up to ~6 hours of total node unavailability per restart rather than one hour.What is now unresolved: why the spawn's own 1-hour timer did not fire, given
DEP0190provesspawn()was reached and the timer is armed immediately afterward. Two candidates, neither confirmed:- That spawn settled and something later is blocking — but the direct
shchild remains an unreaped zombie, which argues against clean settlement. - Preparation is waiting to acquire a lock whose recorded owner is this same process/thread. The observed lock ticket was
{"pid":1,"threadId":0,...}andisOwnerAlive(:1300) isowner.pid !== process.pid || isThreadRunning(owner.threadId)— for a ticket owned by the current pid it reduces toisThreadRunning(0), so a live main thread reads as a live holder. If the waiter is that same thread, it waits on itself until the ~6h budget expires.
Distinguishing these needs an inspector session on the wedged process, which I did not have access to. I would rather leave the mechanism explicitly open than assert the wrong one twice.
Unchanged and still verified: startup is gated on component installation (
server/loadRootComponents.js:18,server/threads/threadServer.js:180-181beforelistenOnPorts()at 225); a component with no lock entry and no installed directory re-enters preparation on every boot (:1389-1391); the node is fully unavailable and logs nothing for the entire window. The headline defect — an unresolvable component package silently taking a whole node out of service on restart — stands; only my timing claim was wrong.Mechanism resolved via inspector session on the wedged process. It is unbounded — neither of my earlier timing claims was right, and the real cause is a self-owned preparation lock whose wait deadline renews forever.
Runtime evidence (read-only introspection, main thread, pid 1, uptime 5431s)
threadsglobal is an empty array — no worker threads started;serverhas no listener. Confirms the block is insideloadRootComponents→installApplications().process._getActiveRequests()→ empty. No pending I/O request.- libuv handle census:
{async:9, timer:2, check:2, idle:1, prepare:1, pipe:2, signal:3, fs_event:2, loop:1}— noprocesshandle, so no live child process. - The two pipes are the process's own
fd:1/fd:2(stdio: "stdout"/"stderr"), not a child's. The "grandchild holds the child's pipes" theory is therefore wrong. - Exactly one persistent timer, same address across samples, re-arming at ~16-50 ms. No long timer exists — no 1-hour spawn timer, no multi-hour lock timer.
Why no timeout will ever fire
components/componentPreparationLock.ts:LOCK_POLL_INTERVAL_MS = 50(:13) andawait delay(LOCK_POLL_INTERVAL_MS)(:295) — matches the observed ~50 ms timer exactly.- The wait bound is a
performance.now()deadline (:251), not asetTimeout— which is why no long timer is armed. - :283-293 — on deadline expiry, if the holder's liveness is confirmed the deadline is renewed:
if (performance.now() >= deadline) { if (await ownerLivenessConfirmed(blocker, options)) { deadline = performance.now() + timeoutMs; // renewed } else { throw new Error(`Timed out waiting for component preparation lock ...`); } }
ownerLivenessConfirmeddelegates tooptions.isOwnerAlive, which atcomponents/Application.ts:1300is:
isOwnerAlive: (owner) => owner.pid !== process.pid || isThreadRunning(owner.threadId)
The observed lock ticket was
{"pid":1,"threadId":0,...}, in a process whereprocess.pid === 1. Soowner.pid !== process.pidis false and it reduces toisThreadRunning(0)— the main thread, which is running, because it is the thread doing the waiting. Liveness is "confirmed", the deadline is renewed, and the loop polls every 50 ms forever.A thread cannot be both the holder and the waiter, but
isOwnerAlivecannot distinguish those cases, so a stale ticket recorded against the current pid/thread produces a permanent, silent, self-inflicted deadlock. The existing comment at :89-92 shows the renew-forever hazard was anticipated; the stricter check still delegates to the same predicate, so it does not catch this case.Also
waitingReported(:279-281) firesonWaitonce, so there is no repeating log to reveal the wait.Corrections to this issue
- "blocks for the full 1-hour spawn timeout" — wrong; that timer is never armed.
- My later "~6 hours via the preparation-lock budget" — also wrong; the budget is renewed rather than enforced.
- Correct statement: startup blocks indefinitely. The observed node passed 90 minutes with no timer that could ever release it, and would not have recovered without intervention.
Suggested fix
isOwnerAliveshould return false when the recorded owner is the current process and the current thread — that combination is impossible for a genuine live holder and means the ticket is stale. Independently, deadline renewal should be bounded (renew at most N times, or cap total wait) so no liveness predicate can produce an unbounded wait.Retitling suggestion: the "1-hour spawn timeout" framing in the title is inaccurate; this is an indefinite startup deadlock on a self-owned component-preparation lock.
- changed the title
[-]A component package that cannot install blocks startup for the full 1-hour spawn timeout with no log output; node reports 'Harper is not running' throughout[/-][+]Startup deadlocks indefinitely waiting on a self-owned component-preparation lock (wait deadline renewed forever); one unresolvable component package takes the whole node down[/+]on Aug 4, 2026 Exposure window — this code is days old, which is likely why it has not been hit before.
components/componentPreparationLock.tsis absent from every OSS release tag checked — v5.1.19, v5.1.23, v5.1.24, v5.1.26, and v5.2.0. It exists only onmain, introduced by5c6b68471"Fix concurrent component installation" (the concurrent-install/node_modulescorruption work also referenced from #1996).It nevertheless is running in the field: verified inside the affected container that
@harperfast/harper-pro@5.2.0bundles a core snapshot containing the file, dated Aug 1 02:21, with the relevant lines at the same offsets asmain(LOCK_POLL_INTERVAL_MS = 50:13, deadline :251, renewal :284-287, poll :295). So the pro build was cut from a core commit later than the OSSv5.2.0tag. Anyone reproducing from the OSS tag will not find this file — check the bundled core in the running package instead.Why the trigger is narrow, which compounds the low exposure:
- Only package-based components are considered at all —
Application.ts:588skips config entries without'package' in applicationConfig. Payload-deployed components (the usualdeploy_componentpath) never enterinstallApplications. - A component that has installed successfully once has both a directory and a matching lock entry, so :1389-1391 skips it on every subsequent boot — it never touches the preparation lock again.
- The vulnerable window is therefore narrow: a package-based component that has never successfully installed, on a boot. With a working package URL that resolves on first boot and is skipped thereafter.
- It additionally needs the package to be permanently unresolvable (here: an SSH remote with no credentials in the container), and restarts are infrequent — the bad entry sat inert until an upgrade forced one.
Net: affected builds are recent pro releases carrying a post-
v5.2.0core snapshot, and the trigger requires a never-installed package-based component plus a restart. That is consistent with this being an early sighting rather than a long-standing silent problem — but it also means exposure grows as that core snapshot propagates into more releases.- Only package-based components are considered at all —
Correction to the "no credentials" claim in the Observed section (body updated). My original wording said the container had no credentials, evidenced by
~/.sshbeing absent — that check was run viadocker execas root, so~expanded to/root/.ssh, which is not where Harper's runtime user lives. Wrong path, and the blanket claim was too strong.Re-checked properly. The corrected position:
- No
.sshexists for either the container root or theharperdbuser, so there is genuinely no SSH key,known_hosts, or host-alias config. - There is a git credential mechanism —
gitCredentialHelper.js/gitCredentialServer.ts, viaHARPER_GIT_CREDENTIAL_SOCKET, wired ascredential.helperorGIT_ASKPASS. But its own comments show it answers HTTPS prompts (Username for 'https://github.com'), i.e. token/password auth — not SSH key auth, so it cannot serve an SSH remote. getGitSSHCommand(which setsGIT_SSH_COMMANDonmain) is not in this build.- Credentials are supplied per deploy request; a config-declared
package:installed at boot has no request to carry them. gititself is present (2.47.3), so the failure is authentication, not a missing binary.
So the conclusion stands — that URL could not have authenticated — but for a more specific reason than originally stated: the instance has an HTTPS credential channel and the component was configured with an SSH remote.
One check I am explicitly withdrawing: I attempted to read
/proc/1/environto enumerate liveSSH_*/GIT_*variables and it failed with permission denied. Any inference that no such variables are set is unsupported — I could not read the process environment.None of this affects the deadlock mechanism, which is in lock acquisition and independent of why the install could not succeed.
- No
Retracting the credentials claim entirely — it was wrong. SSH deploy-key auth was configured on the affected node.
I twice asserted the instance had no way to authenticate the SSH remote. Both checks looked at
~/.ssh, which is not a path Harper uses. The mechanism ismaterializeGitSSH()(components/Application.ts~1326-1345), which readsjoin(getConfigValue(CONFIG_PARAMS.ROOTPATH), 'ssh')— a durable<rootPath>/sshdirectory whose*.keyfiles are materialized into a transient 0700 dir per git-over-SSH spawn, with the ssh config copied and itsIdentityFilelines repointed at the transient copies.On the affected node that directory contained:
- a
configdefining exactly the host alias used in the failing package URL, withIdentityFilepointing at - a present
0600OpenSSH private key (plaintext, i.e. a legacy unsealed key — explicitly supported per the function's own docs: "Legacy plaintext keys … are copied through unchanged"), and - a populated
known_hosts.
So the trigger was not missing credentials. I have removed that claim from the body. The reason the install did not complete is now undetermined and, importantly, is not load-bearing for this issue — the deadlock is in lock acquisition and is independent of why the install failed.
One candidate a maintainer may want to look at, explicitly unproven:
materializeGitSSH()returnsundefinedwhenROOTPATHis unset, its comment noting "config not initialized (e.g. an install-time spawn) — no ssh dir to read". If a boot-timeinstallApplications()spawn runs before that config value is available, noGIT_SSH_COMMANDis exported and git falls back to a~/.sshthat does not exist on this image — an auth failure despite correctly configured keys. That would be a distinct bug.Also correcting two smaller errors from my earlier comment:
getGitSSHCommandis not a symbol in either the running build ormain(it came from an out-of-date local tree), and both trees do containGIT_SSH_COMMANDhandling — so my statement that this build has no SSH-command path was false. The HTTPS credential-helper machinery (gitCredentialHelper.js/gitCredentialServer.ts) is real but simply orthogonal to SSH remotes.Apologies for the churn on the trigger description. The mechanism section — verified against the running bundled core and against runtime inspector state — is unaffected.
- a
- changed the title
[-]Startup deadlocks indefinitely waiting on a self-owned component-preparation lock (wait deadline renewed forever); one unresolvable component package takes the whole node down[/-][+]One component's hung install cleanup deadlocks startup indefinitely — no listener opens, nothing is logged, no timeout can release it (root cause: #2076)[/+]on Aug 4, 2026 Mechanism corrected (third revision) and root cause split out to #2076.
An outside review refuted the self-owned-lock explanation, and it was right.
scanLiveClaimsis invoked withowner.tokenfrom inside the wait loop and filtersclaim.owner?.token === ownToken, so a process cannot block on its own claim —blockerwould be falsy and the loop would break immediately. The unreleased ticket is therefore a symptom of hanging inside the critical section, not the cause. That retraction is now in the body.Following the evidence instead of the hypothesis produced a mechanism that accounts for every observation: the
'close'handler awaitsterminateProcessTree, whose SIGKILL escalation path callswaitForConfirmedTermination(() => processGroupIsAlive(pgid)).processGroupIsAliveprobes withprocess.kill(-pgid, 0), which succeeds for an exited-but-unreaped process, andwaitForConfirmedTerminationhas no deadline (unlike the bounded siblingwaitForProcessGroupExit(pgid, timeoutMs)). With a zombie child in the group, the poll runs everyPROCESS_TERMINATION_POLL_MS = 25forever.That matches the measured state exactly — one ~16-50 ms re-arming timer and no long timer, no
processhandle (so'close'had fired), a zombie child as the thing being waited on, the lock ticket held because we are inside the critical section, sleeping threads at 0.76% CPU, and no output.The primitive is filed on its own as #2076, since it can hang any caller of
terminateProcessTree, not just component installs. This issue now covers what makes it fatal: startup is gated on component preparation, so one component takes the whole node down silently and permanently on any restart.Also settled: the
ROOTPATH-unset hypothesis for the git failure is unreachable —installApplications()initialises config before any spawn — so the trigger for the failed install remains undetermined, and deliberately so; the deadlock does not depend on it.Three superseded explanations are listed explicitly in the body so they are not re-tried.
Both open questions from this issue are now resolved, while the node was still wedged.
1. Why the child was never reaped — answered in #2076. Short version:
detached: truemakes the direct child a group leader; it spawned a grandchild; the direct child exited and was reaped normally (that is what fired'close'); the grandchild was orphaned and re-parented to pid 1, which is node itself in a container with notini/dumb-init. Node only waits on children it spawned via a libuv handle, and pid 1'sSigCgtdoes not include SIGCHLD, so the orphan is never reaped. Its group's only remaining member is that zombie, which is exactly what keepskill(-pgid, 0)succeeding forever. So the precondition for #2076 is structural for containerized deployments, not a rare accident.2. Why the install "failed" — the premise appears to be wrong, and I am withdrawing it. There is no evidence the install failed:
- The durable ssh config was complete and correct:
HostName github.com,User git,IdentityFile <key>,IdentitiesOnly yes, with matchingknown_hostsentries. (An earlier idea that the host alias might not resolve is void —HostNameis set, so the alias never needs DNS.) materializeGitSSH()demonstrably ran: its transient 0700 dir was still on disk with the materialized0600key and rewritten config, mtime matching the spawn to the millisecond. So a workingGIT_SSH_COMMANDwas exported.- Whether the git command itself succeeded or failed is unknowable from outside the process:
nonInteractiveSpawnbuffers stdout/stderr and only prints them once the promise settles, which never happened.
So the observable defect is the cleanup hang, not a clone failure. The component directory was never created, but that is equally consistent with a clone that completed into staging and then hung before being moved into place.
Body updated accordingly: this issue no longer asserts the install failed, only that preparation was entered and never completed — which is all the report needs, since the deadlock is in spawn cleanup and is independent of the git outcome.
One incidental finding worth noting for whoever fixes this: the transient decrypted-key directory outlives its documented lifetime while the spawn is hung (details in #2076). For nodes using sealed
enc:v1:keys that turns an at-rest-encrypted key into an indefinitely-persisted plaintext file.- The durable ssh config was complete and correct:
Second independent field occurrence, on 5.2.13 — reproduced three times on one node inside an hour, with two different package URLs. The mechanism section here holds exactly.
Different cluster, different customer, ~7 weeks after this was filed. I hit this cold during an incident investigation and arrived at the same place from the operator side before finding this issue. Posting the independent confirmation plus four things that are new.
Everything below is from the live node. Cluster, host, component and repository identifiers omitted, matching the convention in the body.
Confirmation — identical signature on
@harperfast/harper-pro@5.2.13Two nodes, node A and node B. Node A entered preparation on boot for a package-based component with no lock entry and no installed directory, and wedged:
harper get_status -> "Harper is not running." openssl s_client 127.0.0.1:9926 (inside the container) -> ECONNREFUSED openssl s_client 127.0.0.1:9933 (inside the container) -> ECONNREFUSED <rootPath>/operations-server -> does not existLast line ever emitted, matching the body verbatim:
(node:1) [DEP0190] DeprecationWarning: Passing args to a child process with shell option true ...Process table, ~19 minutes in:
UID PID PPID C STIME TTY TIME CMD harper 1 0 0 13:52 ? 00:00:04 node /home/harperdb/.npm-global/bin/harper harper 281 1 0 13:53 ? 00:00:00 [git] <defunct>CPU 0.24%, RSS flat at 225 MiB, 28 threads with 1 running / 26 sleeping / 1 zombie. One unreleased preparation ticket, holder already dead:
{"pid":1,"threadId":0,"processInstanceId":"d3a5dfdd-…","token":"ffc37f83-…","ticket":1}So: zombie in the group, no listener, no output, flat CPU, ticket held. The
terminateProcessTree→waitForConfirmedTermination→processGroupIsAlivechain in the body accounts for all of it. One small difference from the original report — here the direct child left as a zombie isgititself rather than an intermediatesh, and it is parented to pid 1 (node, no init). Same structural precondition as described in the #2076 cross-post; it does not need a grandchild to reproduce.components/.deploy-aside/was empty, so there was also no prior version to fall back to.Why this still reproduced: #2085 fixed the predicate on
mainonlyThe obvious question is why this recurred at all when #2085 ("avoid startup hangs on zombie process groups") merged 2026-08-27. Answer: it is on
mainonly and was never cherry-picked tov5.2.ref version isProcessGroupAliveincomponents/Application.tsmain5.3.0-beta.2 present (import + call) v5.25.2.13 absent Confirmed a second way, by reading the bundled core actually running on the wedged node (
@harperfast/harper-pro@5.2.13, currentv5.2head). Both the TypeScript source and the compileddiststill carry the pre-fix body verbatim:// dist/core/components/Application.js:1649 — as shipped in 5.2.13 function processGroupIsAlive(processGroupId) { try { process.kill(-processGroupId, 0); return true; } catch (error) { return error.code === 'EPERM'; } }
isProcessGroupAlivedoes not exist anywhere in that build'smanageThreads.js. I could not find a cherry-pick PR of #2085 againstv5.2.So every 5.2.x node — including the newest, 5.2.13 — still counts an unreaped zombie as a live group member, and 5.3.0-beta.2 is the only release line carrying the fix.
Would #2085 have prevented this?
Reading the merged diff: yes, for this failure mode — though via the predicate rather than a timeout, which the PR description makes easy to miss. The description says
waitForConfirmedTermination"is untouched and stays deliberately unbounded after SIGKILL", which is true of the loop, but the same PR rewired the predicate that loop polls:function processGroupIsAlive(processGroupId: number): boolean { - try { - process.kill(-processGroupId, 0); - return true; - } catch (error: any) { - return error.code === 'EPERM'; - } + return isProcessGroupAlive(processGroupId); }and on
mainisProcessGroupAlivefalls through to a/procscan when the leader is a zombie, requires two consecutive all-zombie snapshots, then returnsfalse. For our group — whose only remaining member is a zombiegit— the predicate flips tofalse, the unbounded loop exits,terminateProcessTreeresolves, and boot proceeds. That is #2076's root cause fixed at the source, so the unbounded wait stops mattering for the zombie case.But it does not close this issue
#2085 removes the trigger observed here; it does not remove the fatality. After it:
waitForConfirmedTerminationis still unbounded by design, so a group with a genuinely live descendant still hangs forever — and still hangs boot.- Startup is still gated on component preparation (Expected [WIP] Enable Node.js Type Stripping #2 untouched).
- Preparation is still unbounded independently (Expected Commit changes to package-lock.json from running npm install #3).
- There is still no progress logging (Expected Default to attempting to serve index.html for a path #4) and no memory of a failed install (Expected Add a CODEOWNERS file to set PR reviewers #5).
So the failure class — one component silently and permanently taking a node down on restart — survives #2085 intact; only this particular path into it is closed, and only on 5.3. That seems worth keeping this issue open for on its own terms.
Ask
A cherry-pick of #2085 to
v5.2would stop this recurring on the 5.2 line, where the deployed fleet actually is. Happy to open it if that is wanted.New 1 — the git outcome really is irrelevant, demonstrated by substitution
The body withdrew the claim that the install failed and called the trigger undetermined. This run supports that directly, because I got to watch two different package identifiers wedge the same way.
The first attempt used a genuinely malformed identifier — a GitHub web tree URL (
https://github.com/<org>/<repo>/tree/<branch>), which is not an installable spec at all — with no credentials, on a node that additionally logged:[main/0] [error]: SSH key <name>.key is encrypted but no secret custody is registered on this node; skipping itand carried
secretCustody: {}. That is about as unambiguously doomed as a clone gets, and I initially wrote it up as the cause.The operator then redeployed with a correct spec —
git+https://github.com/<org>/<repo>.git#<branch>— and the node wedged identically: same[git] <defunct>, same terminalDEP0190line, same absent UDS, same flat CPU. Whatever the fix is, it cannot be "validate the package identifier": a well-formed one deadlocks the same way. Worth keeping in mind against any temptation to close this with input validation.New 2 — why peers cannot tell this from a network fault, mechanically
The body notes peers report
Client network socket disconnected before secure TLS connectand loop their reconciler. The reason is structural and worth recording, because it actively misdirects the on-call.On these hosts nginx listens on
0.0.0.0:9933and proxies to the container's published replication port. nginx accepts the peer's TCP connection whether or not anything is behind it. So from node B:CONNECTED(00000003) no peer certificate available SSL handshake has read 0 bytes and written 358 bytes error:0A000126:SSL routines:ssl3_read_n:unexpected eof while reading— a TCP accept followed by a close before any TLS bytes come back. Meanwhile the same port probed from inside the wedged container is plain
ECONNREFUSED.The peer therefore never sees a refusal, only a half-open-then-dropped connection, which reads as a TLS/CA problem or a flaky link. On this incident that cost real time: the symptom looks a lot like the peer-CA trust class of bug, and the surviving node's logs are full of
Reconciling N wedged subscription(s) … lastCloseCode: 1006climbing forever. Useful triage rule for whoever writes the runbook: probe the port from inside the target's own container. Refused there while the peer reports a pre-TLS drop means "no listener at all / boot never completed", not a TLS problem.New 3 —
hdb_deploymentis an already-wired progress surface that this leaves permanently stuckComing in via
deploy_componentrather than a hand-edited config, there is a system-table row per attempt, and it shows the hang from the outside with no inspector needed:project: <name> package_identifier: <url> status: pending phase: prepare event_log: [ { t: <boot+Ns>, event: "phase", data: { phase: "prepare", status: "start" } } ] started_at: <t> completed_at: null error: nullExactly one event. No second event, no error, no completion — 45+ minutes and counting, and the row replicates cluster-wide, so every node advertises a permanently pending deployment. After the redeploy there were two such rows stuck side by side.
This seems directly relevant to Expected #4 and #5. The surface already exists and is already replicated; it just never gets a
prepareheartbeat, a deadline, or a terminal state. A boundedpreparethat writesstatus: failedhere would satisfy "log progress" and "record failed installs so a known-unresolvable package does not re-block every subsequent boot" without inventing a new mechanism. It would also have made this diagnosable from the API in seconds.Related:
credentials: Noneandrestart_mode: rollingare both on the row, so the deploy path has the metadata to decide whether a given install can possibly succeed before it gates boot on it.New 4 — what separates a wedged node from a healthy one is the deploy method, not the node
Correcting an earlier draft of this comment: I first attributed the asymmetry to node A being a fresh clone of node B. That is wrong, and the real answer isolates the trigger more usefully.
The two nodes carry byte-identical SSH material — same
<rootPath>/ssh/config, same sealed<rootPath>/ssh/*.key(matching sha256) — and both are correctly provisioned with the cluster's shared secret-custody keypair. Node A is not misconfigured.It nonetheless logs, during the boot-time install:
[main/0] [error]: SSH key <name>.key is encrypted but no secret custody is registered on this node; skipping itThis is an ordering bug, and it is reachable on any node.
server/loadRootComponents.js:async function loadRootComponents(isWorkerThread = false) { try { if (isMainThread && !process.env.HARPER_SAFE_MODE) await installApplications(); // line 18 } catch (error) { ... } // ... await loadComponentDirectories(loadedComponents, resources); // line 36
installApplications()is the first statement. Secret custody is itself a built-in component (secretCustody=@/dist/security/keyCustody.jsin the built-ins list), and itsstartOnMainThread— which callsregisterSecretCustody()— does not run untilloadComponentDirectories()at line 36. So for the entire duration of a boot-time component install,getSecretDecryptor()returnsundefinedand every sealedenc:v1:key is skipped, no matter how correctly custody is provisioned. A runtime deploy, after components are loaded, has custody available; a boot-time install never does.The healthy node never logs this only because its components are already installed (lock entry + directory present), so no install spawn runs at boot.
This is worth flagging because the same shape of hypothesis appears in the comment history and was cleared:
materializeGitSSH()returningundefinedwhenROOTPATHis unset during an install-time spawn, ruled unreachable becauseinstallApplications()initialises config before any spawn. That reasoning is correct forROOTPATHand does not carry over to custody, which is registered by a component rather than by config init, and therefore really is absent at that point.It is, however, not what wedged node A here — both failing deploys used HTTPS remotes, and
materializeGitSSH()runs unconditionally before every git spawn regardless of protocol, so the skipped SSH key is logged but never consulted. It is a red herring that happens to log immediately before the spawn warning. Recording it because this issue's history shows how easily this path gets misattributed.What actually separates them is how each component arrived. Four deployment rows on this cluster:
project package_identifierpayload_sizestatus app (working) none 189,869,451 success status-checkgit+https://…/status-check-fabric.git#semver:v1.0.0(public)none success devattempt 1web tree URL (private) none pending / prepare devattempt 2git+https://….git#develop(private)none pending / prepare - Payload deploys never enter this path at all. The working application was uploaded as a ~181 MB payload with no package identifier — no git spawn, so no process group, so nothing to poll. Immune by construction.
- A package deploy against a public repo also succeeds —
status-checkinstalls from a git URL on these same nodes with the same unusable key, because it needs no credentials. - Both wedges are package deploys against a private repo with
credentials: None, and both originated on node A.
So the precondition is narrow and clear: a package-identifier deploy whose git invocation cannot authenticate. Not clone state, not node identity, not a missing key. Node B has the same sealed-and-undecryptable key and the same absent custody; by inspection it would wedge the same way if the same deploy were issued against it (not tested — it is a customer node).
One clone-related observation does survive: a newly cloned node has an empty
.deploy-aside, so there is no previously-working copy to fall back to when preparation strands the component directory.- The body calls the failure latent because the entry sits inert until the next restart. Worth making explicit who writes that entry, and when: the deploy handler itself does, in
components/operations.js—await configUtils.addConfig(req.project, applicationConfig)— keyed on the deploy request'sprojectname rather than the package name. It is written as part of the deploy, before the component is known to be installable. So a deploy whose preparation never completes still leaves behind a config entry that arms the next boot. I removed the entry by hand and the node booted cleanly — listeners up, UDS present,get_status: Available. A redeploy of the same project four minutes later wrote it straight back, and the next restart wedged again. Manual removal is therefore durable only until the next deploy attempt of that project; every retry re-arms the wedge. That ordering — persist the config entry first, find out whether it can install second — looks like a cheap place to intervene alongside Expected Add a CODEOWNERS file to set PR reviewers #5.
Timeline on one node, all within about 35 minutes: wedged on boot → exited → config entry removed by hand → booted clean and served → redeploy re-added the entry → restart → wedged again,
[git] <defunct>, no UDS,DEP0190as the last line.Endorsing Expected #2
Of the five, item 2 is the one that would have contained this incident entirely. A component that cannot be fetched taking down replication, the operations API, the HTTP listener and the status endpoint — on a node whose other component was healthy and whose peer was fine — turns a bad deploy into a total outage with no surface left to report it. The node also silently stops answering the status check that gates GTM, so it drops out of rotation with no signal as to why.
Happy to run anything else against this cluster while it is still in this state, or to re-run a specific probe if a maintainer wants a particular piece of inspector state captured.
The custody-ordering finding in my comment above is now filed on its own as #2780 — secret custody registers after
installApplications(), so sealedenc:v1:SSH keys are always skipped during a boot-time component install, on any node regardless of provisioning.Filing it separately because it is a distinct defect with its own fix, not a manifestation of this one: it fails the install, and it is this issue's startup gating plus #2076's unbounded poll that turn that failed install into a dead node. It also reaches further than git — anything relying on a sealed secret during boot-time component preparation is affected.
Noted there, and repeating here so it isn't re-tried: this is not the
ROOTPATH-unset hypothesis that was correctly ruled out in this thread. That reasoning holds forROOTPATH(config is initialised before any spawn) but does not carry over to custody, which is registered by a component rather than by config init.
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
A component that enters preparation on boot and whose install spawn leaves an unreaped child makes startup deadlock indefinitely. The node is completely unavailable throughout —
harper cluster_statusreportsHarper is not running, no listener opens, replication peers cannot connect — and nothing is logged after the first second of boot. There is no timeout that can release it.The underlying defect is an unbounded process-group termination poll, filed separately as #2076. This issue covers what makes it fatal rather than merely untidy: startup is gated on component preparation, so one component can take the whole node down, silently and permanently, on any restart.
Verified on
@harperfast/harper-pro@5.2.0as running in production; the same code shape is onmain.Observed
A config declared three package-based components. One had no lock entry in
harper-application-lock.jsonand no installed directory, so it entered preparation on boot. Itspackage:was an SSH git remote.The reason that install did not complete is undetermined, and is not required for this report. SSH deploy-key auth was configured:
<rootPath>/ssh/contained aconfigdefining exactly the host alias used by that URL, with anIdentityFilepointing at a present0600OpenSSH private key, plus a populatedknown_hosts. That is the directorymaterializeGitSSH()consumes (join(getConfigValue(CONFIG_PARAMS.ROOTPATH), 'ssh')), and a legacy plaintext key of that form is explicitly supported.git2.47.3 is installed. So this was not a missing-credentials case. (A hypothesis thatROOTPATHmight be unset during a boot-time spawn was checked and judged unreachable —installApplications()initialises config before any spawn.)After the last log line —
— nothing more was ever emitted. State at +90 minutes:
harper cluster_status→Harper is not running.grep -c 'successfully started'→ 0;grep -c 'timed out after'→ 0S(sleeping); CPU 0.76%, memory flatsh) present as a zombie,PPid: 1, unreapedFrom an inspector session on the wedged process (read-only; main thread, pid 1, uptime 5431 s):
threadsglobal is an empty array — no worker threads started;serverhas no listener. The block is insideloadRootComponents→installApplications().process._getActiveRequests()→ empty.{async:9, timer:2, check:2, idle:1, prepare:1, pipe:2, signal:3, fs_event:2, loop:1}— noprocesshandle, so theChildProcesshandle had already closed.fd:1/fd:2, not a child's.Mechanism
The install spawn's
'close'handler awaitsterminateProcessTree, which on the SIGKILL escalation path calls:and in
components/Application.ts:A process that has exited but not been reaped still occupies a PID, so
kill(-pgid, 0)succeeds andprocessGroupIsAlivereturnstruefor a group whose only remaining member is a zombie.SIGTERM/SIGKILLare no-ops against an already-dead process, andwaitForConfirmedTerminationhas no deadline — unlike its siblingwaitForProcessGroupExit(processGroupId, timeoutMs), which is bounded. So the await never completes.This accounts for every observation: a ~25 ms re-arming poll and no long timer; the zombie child being precisely what the loop waits on; no
processhandle, because'close'already fired and closed it; the lock ticket unreleased because the process is still inside the critical section that holds it; sleeping threads at ~0.76% CPU; and no output, indefinitely.Full analysis of the primitive is in #2076.
Why the install path is entered at all
components/Application.tsskips installation only when all three hold:A component that has never successfully installed satisfies none of them, so preparation re-runs on every boot. There is no "this failed before, don't gate startup on it again" state. (An empty directory satisfies
existsSync, so a half-finished install reads as installed — a related trap.)Why it takes the whole node down
server/loadRootComponents.js—if (isMainThread && !process.env.HARPER_SAFE_MODE) await installApplications();server/threads/threadServer.js—loadRootComponents(true)is awaited;listenOnPorts()is only reached after it resolves.No listener opens until every component's preparation resolves, so one component's hung cleanup takes the entire node out rather than degrading that component.
Peers cannot distinguish this from a network fault: they report
Client network socket disconnected before secure TLS connectand loop their subscription reconciler indefinitely, because the replication port never opens.It is also latent — the config entry sits inert until the next restart, so the outage surfaces long after the change that caused it and gets attributed to whatever triggered the restart.
Retracted explanations
Recorded so nobody re-treads them:
'close'awaited without'exit', with a grandchild holding inherited pipes." Refuted — the libuv census shows only the process's ownfd:1/fd:2pipes and noprocesshandle, so no child pipes were held open.'close'did fire; the hang is inside its handler.scanLiveClaimsis called withowner.tokenand filtersclaim.owner?.token === ownToken, so a process cannot block on its own claim;blockerwould be falsy and the wait loop would break immediately. The unreleased ticket is a symptom of hanging inside the critical section, not its cause. (Credit to an outside review for catching this.)Expected
Workaround
HARPER_SAFE_MODEskipsinstallApplications()entirely, so a node already wedged this way can be booted with it set. Otherwise the offending component entry must be removed from the config before restarting.Reproducer
package:points at a git remote whose clone will fail, with noinstall:block.Expect: no log output after the shell-spawn warning,
Harper is not runningindefinitely, a ~25 ms poll as the only timer, and a zombie child.Related
Filed from a field incident; cluster, host, component and repository identifiers omitted.