Repository navigation
deploy_component: failed peer-side payload delivery leaves the peer with an empty component dir and the previous build stranded in .deploy-aside (no rollback) #2061
Description
Activity
Companion issue for the payload-blob root cause that triggered this: #2062
Additional finding: the empty component directory also suppresses the self-heal on restart.
components/Application.ts(~585-613) installs fromapplicationConfig.packageunless all three hold:existsSync(application.dirPath), a lock entry exists inharper-application-lock.json, and that entry deep-equals the config entry. The empty directory left behind by the failed deploy satisfiesexistsSync, and the lock still matches the (unchanged) config — so installation is skipped with "already installed with matching configuration", andcomponentLoaderthen pointsloadComponentat the empty directory.Net effect: the node cannot repair itself on restart, because the broken state is indistinguishable from "installed" to the lock check. An emptiness/validity check on
dirPath(rather than mere existence) would let the normal install path recover it.Reproduced, and it demonstrates the bug is direction-agnostic — it just relocates the broken node.
Recovering the un-deployed node by deploying to it directly succeeded locally (that node now has a complete 1.5 GB component directory at the intended version, with its own
.deploy-asidecorrectly emptied). But the same deploy then failed to replicate back to the other node, and left that node in exactly the state described above: an empty component directory, with its previously-installed build (1.5 GB, verified complete —config.yaml,dist, 267node_modulespackages,.env, native addons) stranded undercomponents/.deploy-aside/<name>/1-<ts>-<uuid>/.So the cluster went from "node A un-deployed" to "node B un-deployed" with no net progress. Each attempt to fix it by deploying moves the empty directory to whichever node is on the receiving end.
Two details worth folding into a fix:
- The stranded aside is the cheapest recovery — it is a complete, correct build of the version you wanted. Restoring it is two renames on the same filesystem, no reinstall and no payload transfer. That the product does not do this automatically (or even point at it) is the core gap.
- The window is invisible while it lasts. Neither affected node had restarted, so in both cases the component kept serving from memory and health looked fine. The empty directory only becomes an outage at the next restart — potentially days later, with nothing linking it back to the deploy.
Correction to scope — this appears to be already fixed on
main, and my line references were from a stale tree. Apologies.I verified against
origin/main(24f67d8). Two things change:1.
mainalready restores the aside.components/Application.tsnow wraps extraction in a try/catch that callsrollbackExtractedDirectory(application, asideStagingDir, asidePath)(~762-772), and adds anExtractionTransactionwithcommit()/rollback()plus adeferCommitoption (~786-800) so a later failure can still restore. Because the payload is consumed bypipeline(tarball, gunzip(), extract(application.dirPath))inside that try (~750), a payload-blob read failure — the exact failure in this report — should now roll back rather than leave an empty directory.rollbackExtractedDirectoryis present only onmain:git grep -creturns 0 at bothv5.1.26andv5.2.0, 3 onmain. So the behaviour described in this issue is real for shipped versions and fixed in unreleased code.Revised ask: confirm the above covers the peer-payload-failure path, and consider whether it warrants a backport to 5.2.x — as shipped, 5.1.x/5.2.0 leave a node silently un-deployed with its previous build stranded, recoverable only by hand.
2. Corrected line references (my originals were read from a working tree at 5.1.3):
Claim Correct location on mainaside staging dir constant components/Application.ts:480—export const ASIDE_STAGING_DIR = '.deploy-aside'aside rename components/Application.ts:731-742install-skip condition components/Application.ts:1389-1391The install-skip claim still stands on
main: it testsexistsSync(application.dirPath)plus a matching lock entry, so an empty directory still reads as installed and suppresses the repair install. That part is worth keeping regardless of the rollback fix, since anything that leaves an empty component dir becomes unrecoverable-on-restart.Pinning down the fix and closing my earlier open question.
The fix is commit
12e35262f"Make redeploy preparation transactional" (2026-08-01) — verified onorigin/main, and not inv5.2.0.I had flagged uncertainty about whether the rollback covers failures after extraction (npm install, replication), since
deferCommit/.rollback()appear only inApplication.tsand not inoperations.js. That resolves: the deferred flow is threaded through —extractApplicationis called withdeferCommit = true(~1254), thentransaction.commit()on success (~1267) andtransaction.rollback()on install failure (~1270). So post-extraction failures restore the aside too, not just extraction errors.Net: the failure mode in this report is addressed on
mainand unaddressed in every shipped release up to and including 5.2.0. The remaining question for maintainers is whether12e35262fshould be backported to a 5.2.x patch, given that as shipped a failed peer deploy leaves a node silently un-deployed with recovery only possible by hand.
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
When
deploy_component's peer-side payload delivery fails, the peer is left with an empty component directory and its previously working build stranded incomponents/.deploy-aside/. Nothing restores the aside. The deploy correctly reports failure to the caller, but the peer is silently left with no component on disk — and because the process is not restarted on the failure path, the old component keeps serving from memory. The degradation is invisible until the next restart, at which point the component is simply gone.Observed on 5.1.26, two-node cluster, deploying a component with a ~91 MiB payload (it bundles native binaries).
Observed state after the failed deploy
On the peer:
du -shconfirms it:49Mfor.deploy-aside,0for the component directory.The stranded build was two minor versions behind what the origin had just installed, so the cluster was left split: origin on the new version, peer with an empty directory on disk and the old version live only in memory.
Why this matters
get_deploymentreports that the deploy failed, not that a node was left un-deployed..deploy-aside/<name>/1-<ts>-<uuid>/exists and moving it back by hand.Expected
If the peer-side deploy fails after the aside has been taken, the aside should be restored — the component directory returned to its prior state — before the failure is reported. Failing that, the failure must at minimum surface as "node X is now un-deployed / left in a staged state" rather than only as a replication error, so the degradation is actionable when it happens rather than discovered at the next restart.
Where
components/operations.js—deployComponent; the aside / stage / replicate sequence.components/deploymentRecorder.ts— peer-side payload read; this is the branch that throws and fails the deploy after the aside is already in place.Related
deploy_componenton peers (adjacent, but that path hangs during full-resync; this one fails fast and leaves the aside behind).Filed from a field incident; cluster and host identifiers omitted.