Repository navigation
Staged component deploys: land build-aside-then-swap as a sequence of small PRs #2315
Description
Activity
Step 1 is in progress — #2345 (draft)
Branch:
claude/deploy-stage-swap-step1. Body of the PR carries the full reviewer notes; the short version:deploy_componentnow builds the replacement at.deploy-staging/<deploymentId>/<component>, validates that tree, and only then activates it — live moves aside, candidate takes its place, root config is published last.That closes all three coupled defects at once, because they had one cause (build workspace, serving path and commit boundary were the same object):
- the live tree was moved aside before the replacement existed, so during a deploy the live path held the new release while its dependencies were still installing — requests hit an unrunnable tree, not just an absent one;
- validation ran after the swap committed, so a component that installed cleanly but threw at load went live anyway while the operation returned an error;
- root config was published before the build and never rolled back, so
installApplications()reinstalled a rejected release at the next restart.
Proven rather than asserted: the new integration test holds a deploy open mid-install and samples the live directory throughout. Against
mainit fails withsaw ["2"].Two limits are deliberate and belong to later steps, stated here so they are not rediscovered as bugs:
- Activation is two renames, so the live pathname is briefly absent. A component that opens its own files during a request can still see a gap. Closing it entirely is the successor work listed under Possible successors.
- Validation is still a no-op on the main thread, and the operations API deploys on the main thread — so operator deploys remain unvalidated, exactly as today. Step 1 fixed the order; making validation reachable there is step 2.
Four findings from the pre-push review are carried openly in the PR description rather than fixed: a stale
configBeforesnapshot, incomplete fsync coverage across the two renames, a stale journal left by a compensated activation, and unreliable fail-closed attribution after the swap.- added a commit that references this issue
on Aug 26, 2026 - added a commit that references this issue
on Sep 2, 2026 Step 2 design note — candidate certification in an isolated validator
Posted here so the plan is durable and reviewable before implementation. Four planning rounds ran against it; the framing gate's remaining objection is recorded at the end rather than settled quietly.
Revised after a planning review graded the first draft
better-alternative-exists. The isolation choice
survived; the layer did not. This note adopts the reviewer's deeper framing and narrows the claim to
what this step can actually deliver.The invariant
No candidate may receive activation or recovery authority until that exact candidate has loaded
successfully under its real application identity and effective root configuration, in a non-serving
execution context. Inability to obtain a verdict is failure.Safe mode: stage, do not activate. Round three was right that "skip certification in safe mode" would
mint.completefrom no verdict — false authority of exactly the kind step 1's recovery trusts. Round four
then offered a better answer than my first fix (activate without.complete): stage-only, pending
certification. Safe mode may not execute configured code, so it cannot certify; but safe mode is also
transient, so the candidate can simply wait. The build is staged and the swap does not happen, and the next
ordinary-mode preparation certifies and activates it. That preserves the operator's ability to stage a fix
from safe mode without ever publishing an uncertified tree, which activate-without-.completedid not.This is the one place the exception is deferrable, and that is exactly why it differs from the branch case
below: there, certification can never succeed until scoped storage exists, so "pending" would be permanent
limbo reported as success.Why the first draft was at the wrong layer
validateComponentLoadsExclusivegates its whole body on!isMainThread, and the operations API deploys
on main — so an operator deploy runs no validation and step 1 reordered a no-op. That much was right.What the first draft missed:
validateCandidateis an optional callback onprepareApplication(), and
of its four production call sites onlydeployComponent()supplies one. Meanwhile
activateCandidateApplication()writes.complete, which step 1's recovery treats as authority for a
"build and validation complete" candidate — recovery will roll such a candidate forward after a crash.So the layer that mints the authority does not require the thing the authority asserts. Fixing only
deploy_component's callback would leave three call sites able to produce.completefor an uncertified
tree, which is precisely the one rule, N sites shape that produced most of step 1's defects across 31
review rounds.The requirement goes at the mint site, not at
prepareApplication. A second planning round pointed out
that requiringprepareApplication()to certify still leavesmarkCandidateComplete()and
activateCandidateApplication()exported and able to write.completewith no verdict — so the invariant
would depend onprepareApplicationremembering to certify, recreating the same N-sites shape one frame
down. Instead.completeis gated on module-internal certification state: the module records that a
validator run completed for that exact candidate id, andmarkCandidateCompleterefuses to write without
it. No proof crosses the module boundary, so there is nothing for an external caller to forge or forget.Round three suggested going further and making both mint functions private, noting correctly that each has
exactly one production caller (both incomponents/Application.ts; every other reference is a test, and
I exportedmarkCandidateCompletein step 1 for one). Privatising is not chosen, for one specific reason:
step 1's activation tests callactivateCandidateApplicationdirectly to exercise the two-rename swap and
its crash states in isolation, which is where much of that step's value was proven. Internal gating gets the
same guarantee — an outside caller cannot mint.complete— without giving that up. Activation and
minting authority are separated instead: the export that survives cannot assert validation.What this step guarantees — and what it does not
Guaranteed: within the lifetime of a preparation, no tree reaches
.completeor activation without a
successful load verdict, on any thread, from any of the four call sites.NOT guaranteed — and this is deliberate, tracked as step 3: a
package:deploy's root-config entry is
written before the build (operations.js), and never rolled back. So:v1 is live → v2 is written to root config → v2 fails validation and is discarded → Harper restarts →
installApplications()reads v2 and callsprepareApplication()→ it is certified at that point and
activated.The restart path re-validates, so it cannot publish a broken v2 — but it can publish a v2 the operator
was told had failed. Closing that needs config publication staged with activation, which is step 3 of
#2315. Scope decision taken explicitly: step 2 stays reviewable in a sitting, and the sequence's premise is
small steps. The PR description will state the guarantee in these terms rather than as "unconditional".Isolation: what it does and does not contain
The reviewer was right that the first draft overstated this, so stated precisely:
- Contained: JS heap, module registry, process-global registrations, Scope residue, component status.
A rejected candidate leaves none of it on a serving thread. - NOT contained: databases, the filesystem, the network, child processes, native addons. A candidate
can delete records or write files before it throws, and a native crash can still take the process down.
Two consequences the first draft got wrong:
- The transient-validation apparatus stays. A worker started through
startWorker()joins the ITC
mesh, so a candidate'sserver.registerOperation()announces to main and a request could route to it
between announcement and thread exit. SorunWithDeployValidationGuard, the status sink, and the
Scope/module collection are retained — not deleted as "made unnecessary by isolation". The validator
also runs detached from the serving topology (below), which is what makes them sufficient rather
than merely helpful. process.exit()was a bad argument against reusing a serving worker —workerProcessGuard.tsalready
intercepts it there. The real arguments against that option are shared serving state and hangs.
Spawn: an ephemeral mode, not a new thread type
startWorker()constructs aMessageChannelper connected port, announces the new port to every peer, and
registers the worker for monitoring and automatic restart. Adding a validation thread type would make
validators eligible for acknowledged broadcasts, and on a 64-worker node with concurrent deploys that is
O(deploys × workers) channels plus add/remove traffic, with unrelated broadcasts able to wait on a
validator.So: a separate spawn path, not a mode flag on
startWorker. Round two was right that "share
construction but skip topology" misdescribes it — the MessageChannel fan-out,ADDED_PORT,addPortand
monitoring all run unconditionally, andautoRestart: false(the job precedent) still meshes the worker.
What is shared is the option construction — resource limits,execArgv/preloads, config — factored out
so the two cannot drift; what is not shared is joining the mesh. That matters beyond cost: a meshed
validator'sserver.registerOperationannounces onto main, and ops traffic could route at a dying
validator, which is the hazarddeployValidationState.tsdocuments. Concurrent validators capped.jobRunner.tsis the closest existing precedent (name: 'job',autoRestart: false, plus a
parentPort.postMessagepath for when a non-main thread needs main to spawn) — but jobs are
fire-and-forget, so the verdict protocol is genuinely new and is the part to review hardest.The terminal protocol
One settlement, exactly once, from whichever of these happens first. All of them are failure except the
first:Outcome Verdict { ok: true }messagepass { ok: false, error }messagefail, with the candidate's error synchronous throw from spawn fail (deploy error, retryable — 503 shape, never success) erroreventfail exitwithout a valid verdictfail malformed or duplicate verdict fail message-channel closure fail deadline elapsed fail parent shutting down fail teardown failure fail, preserving the primary validation error The candidate's own error outranks a cleanup error. The worker is terminated and its exit awaited before
its tree is deleted, so a still-running candidate cannot be racing the sweep.worker.terminate()is
already unsafe under Bun in this codebase, so the force-exit path must be Bun-safe rather than assuming
terminate(). Where no safe in-process force-exit exists on a runtime, certification fails closed on that
runtime rather than pretending the deadline is enforceable. Listener exceptions never escape.Deadlines are per caller, and none of them is the ops-API budget. Round two was right that
"operations-API budget" is both underspecified and wrong: a component may legally declare a
handleApplicationtimeout larger thanoperationsApi.network.timeout, so certification would kill a
candidate serving workers would accept; and a wall clock started at the request would already be overdue
after a longnpm install. So: a certification deadline of its own, measured from validator spawn, defaulting
from the component's own load timeout where it declares one — and specified separately for boot
(installApplications) andaddComponent, which have no request budget at all. No fail-open switch: it
would contradict the invariant.Runtime equivalence — and the one place it is deliberately NOT equivalent
Validation must load the candidate roughly the way boot loads an application, or it proves nothing. Boot
suppliesappName,tryRootConfigMount(appName)androotConfigBranchedDatabases(appName); today's
validation supplies none of these beyond a basename. So this step adds one application-loading helper
resolving those inputs for both boot and validation, rather than hand-plumbing a subset into
workerData— hand-plumbing is how the two drift.Branched databases are the exception, and this was a blocker in round two. A branch's location is
"derived only from the application and database names" (resources/branchDatabase.ts:22-27), so a validator
resolving branches the way boot does opens the same durable store the live application is serving from.
A candidate whosehandleApplicationmigrates or deletes rows and then throws would be rejected while v1
keeps serving the mutated branch — the rejected candidate would have gained lasting authority over
production state, which is the invariant this step exists to establish.removeBranchescannot unsay writes
other threads may already have served.Four options:
- Deploy a branch-configured component uncertified — no verdict, and therefore no
.complete—
chosen. It is the same resolution as safe mode above, and for the same reason: the invariant is about
never minting authority without a verdict, not about refusing work. Such a deploy keeps today's
behaviour exactly (it is unvalidated today), gains no false authority, and step 1's contract handles the
crash case — without.complete, recovery rolls back to the committed tree rather than forward. - Refuse to certify and fail the deploy closed — rounds three and four both preferred this, and it is
rejected deliberately: it removes a capability that works today in order to strengthen a guarantee that
has never applied to it. Recorded as the standing disagreement with the planning gate, and carried into
the PR description rather than settled quietly. - Stage-only, pending certification — round four's suggestion, and right for safe mode (above) but
wrong here: certification cannot succeed for these components until (5) exists, so "pending" is
permanent limbo whiledeploy_componentreports success. A deploy that never takes effect is a worse
answer than an honest uncertified one. - Certify with
appNameand mount but omitrootConfigBranchedDatabases— rejected. Loading a
branch-configured application against the base store is not a rehearsal of the real thing and can
itself write to base, so it would mint authority a run did not earn. Certifying against the wrong store
is worse than not certifying. - A validation-scoped branch key that cannot collide with
resolveBranchPath(appName, …)— the right
long-term answer, and what brings these components into the guarantee. New durable-state design; its own
step. - Accept the durable mutation — rejected: it contradicts the invariant.
Consequence, stated rather than buried: the guarantee excludes branch-configured components until (4)
exists — see (5). Nothing they can do today stops working, and nothing they do earns authority — but "no unvalidated
release is published" does not cover them, and the PR description will say so alongside the restart-revival
hole. Both narrowings share one shape worth naming: no verdict means no authority, never no verdict means
no deploy.Approaches considered
- Certification required by
prepareApplication(), verdict from a detached one-shot validator —
proposed. Closes all four call sites; makes.completemean what recovery already assumes. - Isolated validator wired only into
deploy_component's callback — the first draft. Rejected: three
call sites still mint.completeuncertified. - Delegate to a live serving worker — cheapest, reuses reviewed code, but runs candidate code on a
thread serving traffic and keeps every hazard the guard exists for. - Child process — the only option that survives native crashes and
process.exit, but not
runtime-equivalent to Harper's in-process database and resource singletons, and a second bootstrap path.
Held as a successor; the verdict protocol above is reusable unchanged. - Pooled / long-lived validator — amortises spawn cost but carries module and native residue across
deploys, violating clean-heap-per-verdict. - Drop the gate, validate in-process on main — runs app code on the main thread; a candidate that
blocks the loop stalls the node. - Do less: activate, watch worker load status, roll back — rejected on the merits: the invalid
candidate is live in the meantime, and filesystem rollback cannot undo database, network or
process-global effects.
Phase-event compatibility
operations.jsis today the only place that emitsprepare:doneand then theloadphase around
validateComponentLoads. Moving certification inside the preparation would stop those firing, or reorder
them, for every SSE client that keys progress offphase: load— a wire change with no compatibility
option weighed in round one. So the phase stream is preserved: certification emits the sameload
start/done phases from wherever it now runs, andprepare:donekeeps its position relative to them. Worth
an explicit regression test, since the operation's error is otherwise the only authority a client has.Rollout
Main-thread deploys gain a load phase, latency, and new failure modes. Replicated deploys can originate on
a worker, so the nested-validator path needs theparentPorthop rather than a validator spawning a
validator. Safe mode keeps its no-execution behaviour including preloads. Metrics worth having before this
is on by default: spawn latency, validation latency, timeout rate, premature exit, cleanup failure.Standing disagreement with the planning gate
Four planning rounds; the first three each produced findings that changed this design materially (the mint
layer, the branched-database blocker, the deadline basis, the phase stream, the spawn path, safe-mode
authority). Round four's framing verdict remainsbetter-alternative-existson one point only: it holds
that a branch-configured deploy should be refused rather than deployed uncertified, on the grounds that
compatibility cannot override the invariant the change claims to enforce.That is a values judgement, not a factual finding, and it is resolved the other way on purpose: the
invariant is no verdict means no authority, never no verdict means no deploy. Refusing would delete a
working capability to widen a guarantee that has never covered it. The gate's objection is recorded here and
belongs in the PR description under## For the human reviewer, so a reviewer can overrule it rather than
discover it.- Contained: JS heap, module registry, process-global registrations, Scope residue, component status.
- added a commit that references this issue
on Sep 3, 2026 Step 4 is in review — #2531 (draft)
Branch
claude/harper-2315-step-4-dc4c9b; docs companion HarperFast/documentation#668.The premise moved during step 1's review, and the step was re-scoped before implementation.
maindiscards a rejected or failed candidate immediately, and boot recovery removed every owned, journal-less staging directory — so nothing intentionally leaves an unactivated build behind today; only the crash window between.completeand the journal write does. Step 4 therefore lands the bound and the knob as groundwork for step 6, and makes recovery stop destroying the state step 6 will rely on, rather than bounding something that already accumulates.What lands:
deployment.stagingRetention.maxCount(default 5,0keeps none), same coercion shape aspayloadRetention.maxSize.- A dormant build is
.complete+ the owner's tree + no journal. Boot recovery keeps dormant builds and bounds them per component (newest by.completemtime, id tie-break); residue — partial tree, tree already moved live, stale.unsettled— is still removed as before. - Removal is decided only under the owner's preparation lock. Boot catalogues unlocked and takes the lock once per over-bound owner, re-deriving only that owner's catalogued directories under it (three review rounds converged here: per-directory locking at boot let healthy components lose the 250 ms probe to sibling threads; a whole-root re-read under the lock recreated the same contention).
- The deploy path prunes inside the settlement scan it already runs under the lock, before building;
drop_componentreclaims the dropped component's dormant builds.
Open decisions carried in the PR's For the human reviewer: count rather than size;
.completemtime as the order; default 5 (retains crash-window residue that used to be swept); a held lock on an over-bound owner defers the component like the residue branch does; and unowned residue (empty directories from a crash during resolve) is deliberately still out of scope because it cannot be told from a build that has not reached extraction.For step 6: it must pin the build it is about to activate or prune after staging — the deploy-path prune runs before the build with no notion of a build in use — and it decides whether a stage-only mode wants an explicit dormant marker rather than reusing
.complete.16 remaining items
Step 2, re-planned: canary rollout certification
Replaces the isolated-worker validator shelved on 2026-09-14 (#2476). Proposed by @kriszyp on #2893.
Why now. #2881 showed that the in-process load check before the swap cannot give a trustworthy verdict. The operations API deploys on the main thread, which never validated, and peers validated inside workers whose intrinsics
freeze-after-loadhad already frozen, so every candidate with a dependency that extends an intrinsic at load (reflect-metadata) was refused on every peer while the origin ran it. #2893 lands first as the 5.3 fix: under the default lockdown no node load-validates a deploy, andnone,freezeandseskeep the in-process check. That leaves the default with no validation, which is what this step restores.The idea
Certify a release by rolling it out one worker first:
- Keep the previous release restorable (its files and root-config entry), and start one replacement worker on the new release.
- The replacement reports whether the component loaded before it is admitted to traffic, and before the old worker it replaces is retired where overlap is supported.
- On success, continue the rollout, checking each later replacement too. On failure or timeout, stop the rollout, restore the previous release, and remove or replace any worker running the failed candidate.
Why this instead of the isolated-worker validator
The validator was shelved because no host was both serving-equivalent and force-killable. The canary is a real serving worker booted through the normal path, so its load is the boot load (a fresh realm, before the freeze). There is no second bootstrap, no special worker mode, and no separate verdict protocol for a thread that never serves, and it takes no traffic until it reports.
What it needs that does not exist yet
- A restorable previous release. That is step 5: re-home the displaced tree as a dormant artifact, so restoring is step 6's activation path pointed at the previous release.
- Holding a replacement out of traffic until it reports a component-load verdict. Today a worker starts serving once
loadRootComponentssettles, whether or not each component loaded, so the verdict has to come from component status, not from the worker having started. - Platforms without overlap. On Windows, macOS and Bun a replacement only starts after the old worker exits (
platformCanPreStartReplacementinserver/threads/manageThreads.js), so the canary costs capacity there, and a single-worker node is down for the check. - A hung canary under Bun.
terminate()still segfaults there, so a canary that blocks its event loop can only be abandoned, not killed. Where overlap exists, the old workers keep serving meanwhile.
Open questions
restart: falsedeploys start no worker. Canary anyway, or no verdict?- Peers: each node canaries independently, so a failure that happens on only one machine still leaves nodes on different versions. Per-node verdicts could come back through the origin the way staging's
staged: truemarkers do, and we need to decide whether one node's failure rolls the others back. - The deploy's answer would wait for the canary's verdict, so a
restart: truedeploy gets slower by one worker boot. - What counts as loaded: the component's status reaching
healthy, or an explicit readiness signal for components that finish setting up asynchronously.
Carried from #2893 (merged)
- Retire the in-process check in every mode. Deploys no longer fail on cluster peers when a dependency extends a built-in as it loads, such as reflect-metadata #2893 skips it only under the default
freeze-after-load. Undernone,freezeandses, peers still reject a release that throws at load while the origin, which never validates on the main thread, activates it, so a genuinely broken release still splits the cluster there. Undernonethe check can also reject wrongly when the running version already changed a built-in, for example by making a property non-configurable. Once the canary certifies, the in-process check should go in every mode, so origin and peers validate the same way. - Keep the
loadprogress phase. Since Deploys no longer fail on cluster peers when a dependency extends a built-in as it loads, such as reflect-metadata #2893 it no longer fires from workers, and the operations API never emitted it. The canary's verdict should become the deploy'sloadphase, so progress clients see one signal for it. - Test on two nodes. Deploys no longer fail on cluster peers when a dependency extends a built-in as it loads, such as reflect-metadata #2893 proved the peer path on one node, through
server.operation()from a worker. Per-node verdicts and rollback on peers need a two-node test, and harper-pro'sintegrationTests/clustersuite can host it.
🤖 Drafted with Claude Code
- added a commit that references this issue
on Sep 28, 2026 - added a commit that references this issue
on Sep 29, 2026 - added sub-issues
on Oct 2, 2026 Step 2, designed: canary rollout certification
This settles the 2026-09-28 plan and answers its open questions. It went through three cross-model planning rounds (Codex and Gemini). Each round changed the design, and both reviewers' findings from round three are folded in below. It lands as one PR.
The guarantee, and its limit
Per node, for controlled starts only. While a release is pending certification:
- Every worker started that would load it is held, and only one of them is booted at a time: the canary.
- The canary loads the release through the normal boot path, then waits out of traffic.
- If the canary rejects the release, or the process dies before the canary decides, no worker is ever started on that release on that node again. The node either restores the previous release or, if none is kept, fails the component closed.
What it does not cover. Workers already running when the release swaps in keep the code they loaded. Anything they load afterwards still comes from the live directory:
- static files;
- a module first imported after the swap;
- watcher-driven reloads.
That is how every deploy behaves today. A restore puts the files back, but it cannot unload a module an old worker already imported. #2975 (take the node out of rotation for the whole change) or versioned activation paths would close this gap.
Decided
- The origin goes first, for both
restart: trueand'rolling'. Studio deploys with'rolling'. - Each node gets its own verdict, and there is no cross-node rollback. A peer that rejects a release the origin certified restores only itself, and the deploy (or rolling job) names it.
- Later replacements are held and checked too. If one fails, the worker it would have replaced keeps serving and the rollout stops (
workersKeptOnOldCode). Nothing is restored. restart: falsegets no verdict. This is documented.- Pre-swap certification is out of scope. The goal of keeping the node out of rotation while it changes moved to Load-balancer-aware rolling deploys: take each node out of rotation, drain, deploy, rejoin when healthy #2975.
How it works
- The held start. A worker started with a certification instruction boots exactly as any worker does, then holds before
listenOnPorts():- It reports a verdict built from a private outcome the loader records for that component: executed, skipped, pending or failed. Public health status is not used, because a status update can overwrite a failure, and
loadComponent: dev-onlyor safe mode would report "loaded" without running anything. - The verdict also names the deployment id it loaded, read from the tree's provenance marker before and after the load.
- The worker then waits for main to admit it or shut it down.
- It reports a verdict built from a private outcome the loader records for that component: executed, skipped, pending or failed. Public health status is not used, because a status update can overwrite a failure, and
- The start gate, owned by main.
- Who is gated: every HTTP worker start that places a pending component, whatever started it (a restart, a crash auto-restart, the isolated reconcile). The gate lives in
startWorker, so no caller can bypass it. - Canary and waiters: the first such start is the canary. Other starts wait for its decision.
- Before the swap: while a registration is armed but not yet committed, starts that place the component are deferred. A dead requester is resolved from durable evidence under the preparation lock.
- Restarts: replacement restarts on main are serialized.
- Pre-boot: every replacement in a certifying restart pre-boots held while its old worker serves, pool and dedicated alike, on every platform. Admission then follows the platform's handover rule, so a rejection never empties a slot.
- Deadline: the existing 60-second replacement backstop. A timed-out canary is terminated on Node. On Bun it is abandoned, and that component then refuses certification until the process restarts.
- Who is gated: every HTTP worker start that places a pending component, whatever started it (a restart, a crash auto-restart, the isolated reconcile). The gate lives in
- A durable record,
.deploy-staging/<id>/.certification.json.- Written before the swap, under the preparation lock, after main acknowledges the registration. It names the release this activation will keep as its predecessor.
- States:
pending→certified(the fence holds until the rollout ends) → removed. Orrejected→ removed once the restore lands, or kept if there is nothing to restore, in which case the component fails closed on every thread and is never automatically reinstalled. - The fence: while a record is pending or certified, every other preparation and
drop_componentof that component is refused 409. The one exception is an activation of the same id, which joins the in-flight decision. - Retention: the predecessor the record names is pinned.
- At boot, after activation recovery and before
installApplications():- a
pendingrecord left by a dead process is rejected and restored; - a
certifiedrecord is accepted; - a
rejectedrecord has its restore retried; - a record whose release is not live is removed.
- a
- A failed record write keeps that component's starts blocked, stops the candidate, and still settles every caller.
- The restore is step 6's journaled activation of the kept predecessor, guarded so it acts only while the rejected release is still live. Step 5 keeps the rejected tree dormant under its own id, so
deployment_idcan retry it once its cause is fixed. With no predecessor kept, the component fails closed. That covers a first deploy,deployment_stagingRetention_maxCount: 0, and a tree from a boot install,add_componentor before step 5. A rejected first deploy keeps its tree and entry. Re-activating its id re-certifies it. - A deployed tree is not reinstalled over. Today a package deploy's restart re-resolves the package from source, because no deploy writes
harper-application-lock.json. That would have the canary certify different bytes from the ones deployed. Instead,installApplications()skips a component whose live tree's provenance names a deployment whose declared entry equals the effective entry in force. Rejection is checked before any automatic install, independently of that check. - The deploy.
- Origin: prepare, which registers, writes the record and swaps; then wait for the verdict, which is the
loadphase. On rejection the deploy fails and nothing is replicated. Otherwise it replicates while the rest of the origin's rollout continues. restart: truepeers do the same, asking main through theRESTARTmessage. Main replies before it replaces the requesting worker, which it replaces last.'rolling': the origin certifies, then replicates the release as a stage. A job then activates it on each peer in turn (deploy_component { deployment_id, restart: true }), so a peer's release goes live only at its turn. The job visits every peer and reports each one that did not certify.- The response gains
certification: certified | uncertified | unavailable | not-requested. A rejection is a failed deploy whose error carries the failures and what was restored.
- Origin: prepare, which registers, writes the record and swaps; then wait for the verdict, which is the
- The in-process load check is retired in every mode. Under
none,freezeandses, a worker-run deploy that throws at load used to be refused. It now activates unless a certifying restart follows. This is documented and covered by a test. The transient-validation apparatus that only that check used is removed in a follow-up.
Not covered
restart: false.- Isolation flips.
- Failures after a component's load completes.
- Whatever the rejected release did while loading (database writes, files, network, singletons on the canary's worker index).
- Running workers' later reads of the live directory.
- Cross-node rollback.
- Peers on an older build, which restart uncertified. Clusters run one version.
Verification
- Unit: real modules and fixture workers, with barriers at each boundary:
- arming, writing the record, the swap, the decision, and the restore;
- a worker that crashes during the canary;
- overlapping restarts and concurrent deploys;
- skipped loads and safe mode;
- failed writes, and Bun abandonment.
- New integration suite:
- a held canary while the previous release keeps serving;
- rejection and restore;
- a rejected first deploy that fails closed and is not reinstalled;
- the process killed mid-canary;
'rolling'on one node;- a package deploy's restart not reinstalling;
- a lazy import, documenting the stated limit.
- Cluster, a release gate in a harper-pro companion:
- an origin rejection replicates nothing;
- a peer rejection restores only that peer;
- the rolling job continues past a rejecting peer and an unavailable one;
- a single-worker peer still answers.
🤖 Drafted with Claude Code
- added a commit that references this issue
on Oct 6, 2026 All seven steps have merged. The last, step 2 (canary rollout certification), landed in #2981 with its docs in HarperFast/documentation#707, and ships in v5.4. The milestone moves to v5.4 to match.
The two follow-ups still open, #2573 (the 5.2 backport of step 1) and #2659 (
activated_fromon upgraded installs), now sit under the epic #641, so they stay visible after this closes. The later ideas this sequence left are #2975 (take each node out of rotation for its deploy) and #3025 (server.isPrimaryThread).
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Done. All seven steps have merged. The last, step 2 (canary rollout certification, #2981), ships in v5.4. Its open follow-ups, #2573 and #2659, now sit under the epic #641.
Coordination issue for landing staged (build-aside-then-swap) component deploys as a sequence of small, independently mergeable PRs. Replaces the single large effort in #1849 / #2301, both reference-only.
The claim the sequence is carrying: every node installs the release, a replicated deploy reports any node whose install came out different from the origin's, and no node's live tree changes until a candidate exists and can be swapped in as one transaction that includes the root-config entry. A restarting deploy's release is then certified in a canary before it serves, and put back if it fails. Everything else — retention, revert, deploy-from-an-aside — exists to make that transaction reversible after the fact.
flowchart LR S1["1 · Build aside,<br/>then swap"]:::done S4["4 · Retention of<br/>built asides"]:::done S6["6 · Stage now,<br/>activate later"]:::done S3["3 · Root config as an<br/>activation effect"]:::done S5["5 · Keep the displaced tree,<br/>and make an id mean<br/>something after the swap"]:::done S7["7 · Report when a node<br/>installs differently"]:::done S2["2 · Canary rollout<br/>certification"]:::done S1 --> S4 --> S6 --> S5 S1 --> S3 S1 --> S7 S5 --> S2 classDef done fill:#dcfce7,stroke:#16a34a,stroke-width:2px,color:#14532d classDef next fill:#fef3c7,stroke:#d97706,stroke-width:2px,color:#78350f classDef shelved fill:#fee2e2,stroke:#dc2626,stroke-width:2px,stroke-dasharray:5 4,color:#7f1d1dProgress
claude/deploy-worker-validation-step2@9734beaababandoned/claude/build-once-replicate-step7@ed6ceb6d0; cluster test HarperFast/harper-pro#945; docs HarperFast/documentation#698)The steps
1. Build aside, then swap — merged, #2345. The candidate is built under
.deploy-staging/<id>and activated as two renames with journalled recovery. The live tree stops being a construction site. Not gapless: the live pathname is briefly absent. Follow-up #2893 (merged, fixing #2881; docs: HarperFast/documentation#692) stops the in-process load check on workers under the defaultfreeze-after-loadlockdown, where it refused every candidate whose dependency extends an intrinsic at load, so by default no node load-validated a deploy until step 2 landed; #2981 then removed the in-process check altogether. A node still on 5.3.0-beta.2 or beta.3 keeps refusing such a release until it is upgraded; a retry after the upgrade converges.2. Canary rollout certification — merged, #2981 (docs: HarperFast/documentation#707), shipping in v5.4. It replaces the isolated-worker validator, shelved because no host is both serving-equivalent and force-killable, and also the in-process load probe, removed with the apparatus that kept its throwaway load off the live worker. A deploy with
restart: trueor'rolling'certifies its release in a canary: HTTP worker 0's replacement boots on the release and is held out of traffic until it reports whether the component loaded. The rollout goes on only when it did, and each later replacement is held to its own load; one that fails it keeps its old worker and stops the rollout. A rejection, a canary that exits or never answers, or a rollout interrupted before any canary decided restores the previous release (step 5), or fails the component closed when there is none, and the deploy fails with the decision. With'rolling', the origin certifies first, and a job then activates the staged release on each peer in turn, each certifying with a canary of its own. Verdicts are per node, with no cross-node rollback. A deploy that restarts nothing is not test-loaded, and the response'scertificationsays when a restarting deploy could not be certified. Proposed by @kriszyp on #2893; plan. Follow-ups: #2975, taking each node out of rotation first, which also closes the window where a worker not yet replaced imports a file of the new release during the hold; and #3025,server.isPrimaryThread.3. Root config as an effect of the activation transaction — merged, #2789, with follow-ups #2796 and #2801 (docs: HarperFast/documentation#682). A component's root-config entry now changes only as a journaled effect of an activation that committed: it is published durably after the commit rename, under one lock every runtime config writer shares, and recovery replays it on every roll forward before
installApplications()reads the config. A rejected build, a compensated activation and a roll back never touch config, and a payload deploy removespackage,installandcredentials. Both windows step 6 handed over are closed, and thedeployStageActivate.test.jsdivergence is inverted. #2796 madeset_configurationrewrite the file Harper boots from, and retires a lock ticket its owner cannot delete with a release marker. #2801, from kriszyp's review, refuses with a 409 a deploy or drop thatHARPER_CONFIGorHARPER_SET_CONFIGwould undo, removes a dropped component's entry last so a failed drop puts its tree back, and releases a lock ticket a Windows scanner is holding.4. Retention of built asides — merged, #2531.
deployment_stagingRetention_maxCountbounds how many unactivated staged builds a component keeps.5. Keep the displaced tree, and make a deployment id mean something after the swap — merged, #2897 (docs: HarperFast/documentation#696). Every build records its deployment id in
.harper-deployment.jsoninside its own tree, so the rename that makes it live carries its provenance. A described deploy keeps.deploy-staging/<id>after the swap as that id's record, and the tree an activation displaces goes back into its record, dormant under its original id — sodeployment_idreturns to any release retention still keeps, with no new public operation. Kept releases sharedeployment_stagingRetention_maxCountwith staged builds (default 5, each a full installed copy). Activating the id that is already live answers success without a swap, which closes both step 6's unconvergeable retry and step 3's post-commit publish route: a node that switched answers success, a node holding the artifact switches, a node with neither answers 404. Claims are built aside and renamed onto the id, so a crash can no longer burn one. Ruled on the PR: a hand-crafted_deploymentIdnaming an id the node holds is refused 409, a revert publishes the root-config entry that deployment declared,.deploy-stagingpersists after deploys, and a record whose release is gone answers 404.6. Stage a build now, activate it later — merged, #2605 (docs: HarperFast/documentation#670).
activate: falsebuilds without going live;deployment_idmakes that artifact live later, installing nothing, and step 2 certifies it then if that activation restarts. Every deploy returns itsdeployment_id. Mixed-version clusters are a settled trade: the origin detects a pre-5.3 peer afterwards rather than preventing it, because clusters are kept at consistent versions and a second wire path would encumber the common case.7. Report when a node installs something different — merged: #2929, which closed #2295; the follow-up #2954, which records a local
file:source asunidentifiedso a peer never matches on no evidence; the cluster test HarperFast/harper-pro#950; and docs HarperFast/documentation#703.node_modulesincluded, to every peer. It was closed as too heavy, along with its cluster test Every peer runs the origin's build of a replicated deploy (core bump for harper#2315 step 7) harper-pro#945 and docs Document that a replicated deploy runs the origin's build on every node documentation#698. The branches are saved asabandoned/claude/build-once-replicate-step7(@ed6ceb6d0),abandoned/claude/replicated-build-testandabandoned/docs/build-once-replicate-step7.name@version, or npm's own integrity where npm reports no version or commit. A localfile:path isunidentified, since each node reads its own copy (Record a local file: source as unidentified, so a peer never matches on no evidence #2954);warningevent the CLI prints;install_matchesandinstall_differsin each peer'sget_deploymententry.Detail
The long form lives in comments so this stays readable:
main, what is deliberately not in the sequence, unscheduled successors, and the reference PRs.Not in the sequence: #2294 (cross-node activation ordering) and #2273 (a Windows integration failure), both tracked separately.