Skip to content

deploy_component: failed peer-side payload delivery leaves the peer with an empty component dir and the previous build stranded in .deploy-aside (no rollback) #2061

Description

@heskew

Summary

When deploy_component's peer-side payload delivery fails, the peer is left with an empty component directory and its previously working build stranded in components/.deploy-aside/. Nothing restores the aside. The deploy correctly reports failure to the caller, but the peer is silently left with no component on disk — and because the process is not restarted on the failure path, the old component keeps serving from memory. The degradation is invisible until the next restart, at which point the component is simply gone.

Observed on 5.1.26, two-node cluster, deploying a component with a ~91 MiB payload (it bundles native binaries).

Observed state after the failed deploy

On the peer:

components/<name>/                              <- empty (0 bytes), created at the moment of failure
components/.deploy-aside/<name>/1-<ts>-<uuid>/   <- previous working build, complete
                                                    (.env, config.yaml, dist/, docs/,
                                                     node_modules/, package-lock.json,
                                                     package.json, schemas/, templates/)

du -sh confirms it: 49M for .deploy-aside, 0 for the component directory.

The stranded build was two minor versions behind what the origin had just installed, so the cluster was left split: origin on the new version, peer with an empty directory on disk and the old version live only in memory.

Why this matters

  • Silent. The operator sees a deploy failure that talks about replication; nothing indicates the peer lost its component directory. get_deployment reports that the deploy failed, not that a node was left un-deployed.
  • Delayed blast radius. The component survives in memory, so health checks and traffic look normal. The failure materializes on an unrelated future restart, far in time from the deploy that caused it — and by then the aside directory is the only copy, with nothing pointing an operator at it.
  • Recovery is manual and undocumented. Restoring means knowing that .deploy-aside/<name>/1-<ts>-<uuid>/ exists and moving it back by hand.

Expected

If the peer-side deploy fails after the aside has been taken, the aside should be restored — the component directory returned to its prior state — before the failure is reported. Failing that, the failure must at minimum surface as "node X is now un-deployed / left in a staged state" rather than only as a replication error, so the degradation is actionable when it happens rather than discovered at the next restart.

Where

  • components/operations.js — deployComponent; the aside / stage / replicate sequence.
  • components/deploymentRecorder.ts — peer-side payload read; this is the branch that throws and fails the deploy after the aside is already in place.

Related

Filed from a field incident; cluster and host identifiers omitted.

Activity

  1. heskew commented on Aug 3, 2026

    @heskew
    ContributorAuthor

    Companion issue for the payload-blob root cause that triggered this: #2062

  2. heskew commented on Aug 3, 2026

    @heskew
    ContributorAuthor

    Additional finding: the empty component directory also suppresses the self-heal on restart.

    components/Application.ts (~585-613) installs from applicationConfig.package unless all three hold: existsSync(application.dirPath), a lock entry exists in harper-application-lock.json, and that entry deep-equals the config entry. The empty directory left behind by the failed deploy satisfies existsSync, and the lock still matches the (unchanged) config — so installation is skipped with "already installed with matching configuration", and componentLoader then points loadComponent at the empty directory.

    Net effect: the node cannot repair itself on restart, because the broken state is indistinguishable from "installed" to the lock check. An emptiness/validity check on dirPath (rather than mere existence) would let the normal install path recover it.

  3. heskew commented on Aug 3, 2026

    @heskew
    ContributorAuthor

    Reproduced, and it demonstrates the bug is direction-agnostic — it just relocates the broken node.

    Recovering the un-deployed node by deploying to it directly succeeded locally (that node now has a complete 1.5 GB component directory at the intended version, with its own .deploy-aside correctly emptied). But the same deploy then failed to replicate back to the other node, and left that node in exactly the state described above: an empty component directory, with its previously-installed build (1.5 GB, verified complete — config.yaml, dist, 267 node_modules packages, .env, native addons) stranded under components/.deploy-aside/<name>/1-<ts>-<uuid>/.

    So the cluster went from "node A un-deployed" to "node B un-deployed" with no net progress. Each attempt to fix it by deploying moves the empty directory to whichever node is on the receiving end.

    Two details worth folding into a fix:

    1. The stranded aside is the cheapest recovery — it is a complete, correct build of the version you wanted. Restoring it is two renames on the same filesystem, no reinstall and no payload transfer. That the product does not do this automatically (or even point at it) is the core gap.
    2. The window is invisible while it lasts. Neither affected node had restarted, so in both cases the component kept serving from memory and health looked fine. The empty directory only becomes an outage at the next restart — potentially days later, with nothing linking it back to the deploy.
  4. self-assigned this
    on Aug 3, 2026
  5. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Correction to scope — this appears to be already fixed on main, and my line references were from a stale tree. Apologies.

    I verified against origin/main (24f67d8). Two things change:

    1. main already restores the aside. components/Application.ts now wraps extraction in a try/catch that calls rollbackExtractedDirectory(application, asideStagingDir, asidePath) (~762-772), and adds an ExtractionTransaction with commit()/rollback() plus a deferCommit option (~786-800) so a later failure can still restore. Because the payload is consumed by pipeline(tarball, gunzip(), extract(application.dirPath)) inside that try (~750), a payload-blob read failure — the exact failure in this report — should now roll back rather than leave an empty directory.

    rollbackExtractedDirectory is present only on main: git grep -c returns 0 at both v5.1.26 and v5.2.0, 3 on main. So the behaviour described in this issue is real for shipped versions and fixed in unreleased code.

    Revised ask: confirm the above covers the peer-payload-failure path, and consider whether it warrants a backport to 5.2.x — as shipped, 5.1.x/5.2.0 leave a node silently un-deployed with its previous build stranded, recoverable only by hand.

    2. Corrected line references (my originals were read from a working tree at 5.1.3):

    Claim Correct location on main
    aside staging dir constant components/Application.ts:480 — export const ASIDE_STAGING_DIR = '.deploy-aside'
    aside rename components/Application.ts:731-742
    install-skip condition components/Application.ts:1389-1391

    The install-skip claim still stands on main: it tests existsSync(application.dirPath) plus a matching lock entry, so an empty directory still reads as installed and suppresses the repair install. That part is worth keeping regardless of the rollback fix, since anything that leaves an empty component dir becomes unrecoverable-on-restart.

  6. heskew commented on Aug 4, 2026

    @heskew
    ContributorAuthor

    Pinning down the fix and closing my earlier open question.

    The fix is commit 12e35262f "Make redeploy preparation transactional" (2026-08-01) — verified on origin/main, and not in v5.2.0.

    I had flagged uncertainty about whether the rollback covers failures after extraction (npm install, replication), since deferCommit/.rollback() appear only in Application.ts and not in operations.js. That resolves: the deferred flow is threaded through — extractApplication is called with deferCommit = true (~1254), then transaction.commit() on success (~1267) and transaction.rollback() on install failure (~1270). So post-extraction failures restore the aside too, not just extraction errors.

    Net: the failure mode in this report is addressed on main and unaddressed in every shipped release up to and including 5.2.0. The remaining question for maintainers is whether 12e35262f should be backported to a 5.2.x patch, given that as shipped a failed peer deploy leaves a node silently un-deployed with recovery only possible by hand.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

P2

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions