Skip to content

Staged component deploys: land build-aside-then-swap as a sequence of small PRs #2315

Description

@dawsontoth

Done. All seven steps have merged. The last, step 2 (canary rollout certification, #2981), ships in v5.4. Its open follow-ups, #2573 and #2659, now sit under the epic #641.

Coordination issue for landing staged (build-aside-then-swap) component deploys as a sequence of small, independently mergeable PRs. Replaces the single large effort in #1849 / #2301, both reference-only.

The claim the sequence is carrying: every node installs the release, a replicated deploy reports any node whose install came out different from the origin's, and no node's live tree changes until a candidate exists and can be swapped in as one transaction that includes the root-config entry. A restarting deploy's release is then certified in a canary before it serves, and put back if it fails. Everything else — retention, revert, deploy-from-an-aside — exists to make that transaction reversible after the fact.

flowchart LR
    S1["1 · Build aside,<br/>then swap"]:::done
    S4["4 · Retention of<br/>built asides"]:::done
    S6["6 · Stage now,<br/>activate later"]:::done
    S3["3 · Root config as an<br/>activation effect"]:::done
    S5["5 · Keep the displaced tree,<br/>and make an id mean<br/>something after the swap"]:::done
    S7["7 · Report when a node<br/>installs differently"]:::done
    S2["2 · Canary rollout<br/>certification"]:::done

    S1 --> S4 --> S6 --> S5
    S1 --> S3
    S1 --> S7
    S5 --> S2

    classDef done fill:#dcfce7,stroke:#16a34a,stroke-width:2px,color:#14532d
    classDef next fill:#fef3c7,stroke:#d97706,stroke-width:2px,color:#78350f
    classDef shelved fill:#fee2e2,stroke:#dc2626,stroke-width:2px,stroke-dasharray:5 4,color:#7f1d1d
Loading

Progress

Step Status PR
1. Build aside, then swap merged #2345, follow-up #2893 (fixes #2881)
2. Canary rollout certification merged, ships in v5.4 (replaced the isolated-worker validator, shelved 2026-09-14) #2981 (docs: HarperFast/documentation#707); the shelved attempt: #2476, branch claude/deploy-worker-validation-step2 @ 9734beaab
3. Root config as an activation effect merged #2789, follow-ups #2796 and #2801
4. Retention of built asides merged #2531
5. Keep the displaced tree, and make an id mean something after the swap merged #2897 (docs: HarperFast/documentation#696)
6. Stage a build now, activate it later merged #2605
7. Report when a node installs something different merged (redesigned twice on 2026-09-29) #2929, follow-up #2954, cluster test HarperFast/harper-pro#950, docs HarperFast/documentation#703 (earlier closed designs: #2917, branch abandoned/claude/build-once-replicate-step7 @ ed6ceb6d0; cluster test HarperFast/harper-pro#945; docs HarperFast/documentation#698)

The steps

1. Build aside, then swap — merged, #2345. The candidate is built under .deploy-staging/<id> and activated as two renames with journalled recovery. The live tree stops being a construction site. Not gapless: the live pathname is briefly absent. Follow-up #2893 (merged, fixing #2881; docs: HarperFast/documentation#692) stops the in-process load check on workers under the default freeze-after-load lockdown, where it refused every candidate whose dependency extends an intrinsic at load, so by default no node load-validated a deploy until step 2 landed; #2981 then removed the in-process check altogether. A node still on 5.3.0-beta.2 or beta.3 keeps refusing such a release until it is upgraded; a retry after the upgrade converges.

2. Canary rollout certification — merged, #2981 (docs: HarperFast/documentation#707), shipping in v5.4. It replaces the isolated-worker validator, shelved because no host is both serving-equivalent and force-killable, and also the in-process load probe, removed with the apparatus that kept its throwaway load off the live worker. A deploy with restart: true or 'rolling' certifies its release in a canary: HTTP worker 0's replacement boots on the release and is held out of traffic until it reports whether the component loaded. The rollout goes on only when it did, and each later replacement is held to its own load; one that fails it keeps its old worker and stops the rollout. A rejection, a canary that exits or never answers, or a rollout interrupted before any canary decided restores the previous release (step 5), or fails the component closed when there is none, and the deploy fails with the decision. With 'rolling', the origin certifies first, and a job then activates the staged release on each peer in turn, each certifying with a canary of its own. Verdicts are per node, with no cross-node rollback. A deploy that restarts nothing is not test-loaded, and the response's certification says when a restarting deploy could not be certified. Proposed by @kriszyp on #2893; plan. Follow-ups: #2975, taking each node out of rotation first, which also closes the window where a worker not yet replaced imports a file of the new release during the hold; and #3025, server.isPrimaryThread.

3. Root config as an effect of the activation transaction — merged, #2789, with follow-ups #2796 and #2801 (docs: HarperFast/documentation#682). A component's root-config entry now changes only as a journaled effect of an activation that committed: it is published durably after the commit rename, under one lock every runtime config writer shares, and recovery replays it on every roll forward before installApplications() reads the config. A rejected build, a compensated activation and a roll back never touch config, and a payload deploy removes package, install and credentials. Both windows step 6 handed over are closed, and the deployStageActivate.test.js divergence is inverted. #2796 made set_configuration rewrite the file Harper boots from, and retires a lock ticket its owner cannot delete with a release marker. #2801, from kriszyp's review, refuses with a 409 a deploy or drop that HARPER_CONFIG or HARPER_SET_CONFIG would undo, removes a dropped component's entry last so a failed drop puts its tree back, and releases a lock ticket a Windows scanner is holding.

4. Retention of built asides — merged, #2531. deployment_stagingRetention_maxCount bounds how many unactivated staged builds a component keeps.

5. Keep the displaced tree, and make a deployment id mean something after the swap — merged, #2897 (docs: HarperFast/documentation#696). Every build records its deployment id in .harper-deployment.json inside its own tree, so the rename that makes it live carries its provenance. A described deploy keeps .deploy-staging/<id> after the swap as that id's record, and the tree an activation displaces goes back into its record, dormant under its original id — so deployment_id returns to any release retention still keeps, with no new public operation. Kept releases share deployment_stagingRetention_maxCount with staged builds (default 5, each a full installed copy). Activating the id that is already live answers success without a swap, which closes both step 6's unconvergeable retry and step 3's post-commit publish route: a node that switched answers success, a node holding the artifact switches, a node with neither answers 404. Claims are built aside and renamed onto the id, so a crash can no longer burn one. Ruled on the PR: a hand-crafted _deploymentId naming an id the node holds is refused 409, a revert publishes the root-config entry that deployment declared, .deploy-staging persists after deploys, and a record whose release is gone answers 404.

6. Stage a build now, activate it later — merged, #2605 (docs: HarperFast/documentation#670). activate: false builds without going live; deployment_id makes that artifact live later, installing nothing, and step 2 certifies it then if that activation restarts. Every deploy returns its deployment_id. Mixed-version clusters are a settled trade: the origin detects a pre-5.3 peer afterwards rather than preventing it, because clusters are kept at consistent versions and a second wire path would encumber the common case.

7. Report when a node installs something different — merged: #2929, which closed #2295; the follow-up #2954, which records a local file: source as unidentified so a peer never matches on no evidence; the cluster test HarperFast/harper-pro#950; and docs HarperFast/documentation#703.

  • Two designs were set aside on 2026-09-29.
  • Decided 2026-09-29: report divergence, don't prevent it. Each node still resolves and installs for itself. It records what it installed:
  • The origin compares every peer with itself, once. A difference shows as:
    • a warning event the CLI prints;
    • a sentence on the deploy's message naming the peer;
    • install_matches and install_differs in each peer's get_deployment entry.
  • A difference never fails the deploy, and a match means equal evidence, not identical trees.
  • Not covered: a restart can still re-resolve a package component, and a node that joins after a deploy installs from configuration.

Detail

The long form lives in comments so this stays readable:

  • Canary rollout plan — step 2 re-planned: the rollout, why it replaces the isolated-worker validator, what it needs from step 5 and the restart path, and its open questions.
  • Per-step detail — why the isolated-worker validator was shelved, what steps 1 and 4 established for everything after them, step 5's evidence, and the table of what step 6 deferred and to where.
  • Background and reference — why this was re-planned, what already exists on main, what is deliberately not in the sequence, unscheduled successors, and the reference PRs.

Not in the sequence: #2294 (cross-node activation ordering) and #2273 (a Windows integration failure), both tracked separately.

Activity

  1. added theissue type on Aug 25, 2026
  2. added this to the v5.3 milestone on Aug 25, 2026
  3. dawsontoth commented on Aug 26, 2026

    @dawsontoth
    ContributorAuthor

    Step 1 is in progress — #2345 (draft)

    Branch: claude/deploy-stage-swap-step1. Body of the PR carries the full reviewer notes; the short version:

    deploy_component now builds the replacement at .deploy-staging/<deploymentId>/<component>, validates that tree, and only then activates it — live moves aside, candidate takes its place, root config is published last.

    That closes all three coupled defects at once, because they had one cause (build workspace, serving path and commit boundary were the same object):

    • the live tree was moved aside before the replacement existed, so during a deploy the live path held the new release while its dependencies were still installing — requests hit an unrunnable tree, not just an absent one;
    • validation ran after the swap committed, so a component that installed cleanly but threw at load went live anyway while the operation returned an error;
    • root config was published before the build and never rolled back, so installApplications() reinstalled a rejected release at the next restart.

    Proven rather than asserted: the new integration test holds a deploy open mid-install and samples the live directory throughout. Against main it fails with saw ["2"].

    Two limits are deliberate and belong to later steps, stated here so they are not rediscovered as bugs:

    1. Activation is two renames, so the live pathname is briefly absent. A component that opens its own files during a request can still see a gap. Closing it entirely is the successor work listed under Possible successors.
    2. Validation is still a no-op on the main thread, and the operations API deploys on the main thread — so operator deploys remain unvalidated, exactly as today. Step 1 fixed the order; making validation reachable there is step 2.

    Four findings from the pre-push review are carried openly in the PR description rather than fixed: a stale configBefore snapshot, incomplete fsync coverage across the two renames, a stale journal left by a compensated activation, and unreliable fail-closed attribution after the swap.

  4. dawsontoth commented on Sep 2, 2026

    @dawsontoth
    ContributorAuthor

    Step 2 design note — candidate certification in an isolated validator

    Posted here so the plan is durable and reviewable before implementation. Four planning rounds ran against it; the framing gate's remaining objection is recorded at the end rather than settled quietly.

    Revised after a planning review graded the first draft better-alternative-exists. The isolation choice
    survived; the layer did not. This note adopts the reviewer's deeper framing and narrows the claim to
    what this step can actually deliver.

    The invariant

    No candidate may receive activation or recovery authority until that exact candidate has loaded
    successfully under its real application identity and effective root configuration, in a non-serving
    execution context. Inability to obtain a verdict is failure.

    Safe mode: stage, do not activate. Round three was right that "skip certification in safe mode" would
    mint .complete from no verdict — false authority of exactly the kind step 1's recovery trusts. Round four
    then offered a better answer than my first fix (activate without .complete): stage-only, pending
    certification
    . Safe mode may not execute configured code, so it cannot certify; but safe mode is also
    transient, so the candidate can simply wait. The build is staged and the swap does not happen, and the next
    ordinary-mode preparation certifies and activates it. That preserves the operator's ability to stage a fix
    from safe mode without ever publishing an uncertified tree, which activate-without-.complete did not.

    This is the one place the exception is deferrable, and that is exactly why it differs from the branch case
    below: there, certification can never succeed until scoped storage exists, so "pending" would be permanent
    limbo reported as success.

    Why the first draft was at the wrong layer

    validateComponentLoadsExclusive gates its whole body on !isMainThread, and the operations API deploys
    on main — so an operator deploy runs no validation and step 1 reordered a no-op. That much was right.

    What the first draft missed: validateCandidate is an optional callback on prepareApplication(), and
    of its four production call sites only deployComponent() supplies one. Meanwhile
    activateCandidateApplication() writes .complete, which step 1's recovery treats as authority for a
    "build and validation complete" candidate — recovery will roll such a candidate forward after a crash.

    So the layer that mints the authority does not require the thing the authority asserts. Fixing only
    deploy_component's callback would leave three call sites able to produce .complete for an uncertified
    tree, which is precisely the one rule, N sites shape that produced most of step 1's defects across 31
    review rounds.

    The requirement goes at the mint site, not at prepareApplication. A second planning round pointed out
    that requiring prepareApplication() to certify still leaves markCandidateComplete() and
    activateCandidateApplication() exported and able to write .complete with no verdict — so the invariant
    would depend on prepareApplication remembering to certify, recreating the same N-sites shape one frame
    down. Instead .complete is gated on module-internal certification state: the module records that a
    validator run completed for that exact candidate id, and markCandidateComplete refuses to write without
    it. No proof crosses the module boundary, so there is nothing for an external caller to forge or forget.

    Round three suggested going further and making both mint functions private, noting correctly that each has
    exactly one production caller (both in components/Application.ts; every other reference is a test, and
    I exported markCandidateComplete in step 1 for one). Privatising is not chosen, for one specific reason:
    step 1's activation tests call activateCandidateApplication directly to exercise the two-rename swap and
    its crash states in isolation, which is where much of that step's value was proven. Internal gating gets the
    same guarantee — an outside caller cannot mint .complete — without giving that up. Activation and
    minting authority are separated instead: the export that survives cannot assert validation.

    What this step guarantees — and what it does not

    Guaranteed: within the lifetime of a preparation, no tree reaches .complete or activation without a
    successful load verdict, on any thread, from any of the four call sites.

    NOT guaranteed — and this is deliberate, tracked as step 3: a package: deploy's root-config entry is
    written before the build (operations.js), and never rolled back. So:

    v1 is live → v2 is written to root config → v2 fails validation and is discarded → Harper restarts →
    installApplications() reads v2 and calls prepareApplication() → it is certified at that point and
    activated.

    The restart path re-validates, so it cannot publish a broken v2 — but it can publish a v2 the operator
    was told had failed. Closing that needs config publication staged with activation, which is step 3 of
    #2315. Scope decision taken explicitly: step 2 stays reviewable in a sitting, and the sequence's premise is
    small steps. The PR description will state the guarantee in these terms rather than as "unconditional".

    Isolation: what it does and does not contain

    The reviewer was right that the first draft overstated this, so stated precisely:

    • Contained: JS heap, module registry, process-global registrations, Scope residue, component status.
      A rejected candidate leaves none of it on a serving thread.
    • NOT contained: databases, the filesystem, the network, child processes, native addons. A candidate
      can delete records or write files before it throws, and a native crash can still take the process down.

    Two consequences the first draft got wrong:

    1. The transient-validation apparatus stays. A worker started through startWorker() joins the ITC
      mesh, so a candidate's server.registerOperation() announces to main and a request could route to it
      between announcement and thread exit. So runWithDeployValidationGuard, the status sink, and the
      Scope/module collection are retained — not deleted as "made unnecessary by isolation". The validator
      also runs detached from the serving topology (below), which is what makes them sufficient rather
      than merely helpful.
    2. process.exit() was a bad argument against reusing a serving worker — workerProcessGuard.ts already
      intercepts it there. The real arguments against that option are shared serving state and hangs.

    Spawn: an ephemeral mode, not a new thread type

    startWorker() constructs a MessageChannel per connected port, announces the new port to every peer, and
    registers the worker for monitoring and automatic restart. Adding a validation thread type would make
    validators eligible for acknowledged broadcasts, and on a 64-worker node with concurrent deploys that is
    O(deploys × workers) channels plus add/remove traffic, with unrelated broadcasts able to wait on a
    validator.

    So: a separate spawn path, not a mode flag on startWorker. Round two was right that "share
    construction but skip topology" misdescribes it — the MessageChannel fan-out, ADDED_PORT, addPort and
    monitoring all run unconditionally, and autoRestart: false (the job precedent) still meshes the worker.
    What is shared is the option construction — resource limits, execArgv/preloads, config — factored out
    so the two cannot drift; what is not shared is joining the mesh. That matters beyond cost: a meshed
    validator's server.registerOperation announces onto main, and ops traffic could route at a dying
    validator, which is the hazard deployValidationState.ts documents. Concurrent validators capped.

    jobRunner.ts is the closest existing precedent (name: 'job', autoRestart: false, plus a
    parentPort.postMessage path for when a non-main thread needs main to spawn) — but jobs are
    fire-and-forget, so the verdict protocol is genuinely new and is the part to review hardest.

    The terminal protocol

    One settlement, exactly once, from whichever of these happens first. All of them are failure except the
    first:

    Outcome Verdict
    { ok: true } message pass
    { ok: false, error } message fail, with the candidate's error
    synchronous throw from spawn fail (deploy error, retryable — 503 shape, never success)
    error event fail
    exit without a valid verdict fail
    malformed or duplicate verdict fail
    message-channel closure fail
    deadline elapsed fail
    parent shutting down fail
    teardown failure fail, preserving the primary validation error

    The candidate's own error outranks a cleanup error. The worker is terminated and its exit awaited before
    its tree is deleted
    , so a still-running candidate cannot be racing the sweep. worker.terminate() is
    already unsafe under Bun in this codebase, so the force-exit path must be Bun-safe rather than assuming
    terminate(). Where no safe in-process force-exit exists on a runtime, certification fails closed on that
    runtime rather than pretending the deadline is enforceable. Listener exceptions never escape.

    Deadlines are per caller, and none of them is the ops-API budget. Round two was right that
    "operations-API budget" is both underspecified and wrong: a component may legally declare a
    handleApplication timeout larger than operationsApi.network.timeout, so certification would kill a
    candidate serving workers would accept; and a wall clock started at the request would already be overdue
    after a long npm install. So: a certification deadline of its own, measured from validator spawn, defaulting
    from the component's own load timeout where it declares one — and specified separately for boot
    (installApplications) and addComponent, which have no request budget at all. No fail-open switch: it
    would contradict the invariant.

    Runtime equivalence — and the one place it is deliberately NOT equivalent

    Validation must load the candidate roughly the way boot loads an application, or it proves nothing. Boot
    supplies appName, tryRootConfigMount(appName) and rootConfigBranchedDatabases(appName); today's
    validation supplies none of these beyond a basename. So this step adds one application-loading helper
    resolving those inputs for both boot and validation
    , rather than hand-plumbing a subset into
    workerData — hand-plumbing is how the two drift.

    Branched databases are the exception, and this was a blocker in round two. A branch's location is
    "derived only from the application and database names" (resources/branchDatabase.ts:22-27), so a validator
    resolving branches the way boot does opens the same durable store the live application is serving from.
    A candidate whose handleApplication migrates or deletes rows and then throws would be rejected while v1
    keeps serving the mutated branch — the rejected candidate would have gained lasting authority over
    production state, which is the invariant this step exists to establish. removeBranches cannot unsay writes
    other threads may already have served.

    Four options:

    1. Deploy a branch-configured component uncertified — no verdict, and therefore no .complete —
      chosen. It is the same resolution as safe mode above, and for the same reason: the invariant is about
      never minting authority without a verdict, not about refusing work. Such a deploy keeps today's
      behaviour exactly (it is unvalidated today), gains no false authority, and step 1's contract handles the
      crash case — without .complete, recovery rolls back to the committed tree rather than forward.
    2. Refuse to certify and fail the deploy closed — rounds three and four both preferred this, and it is
      rejected deliberately: it removes a capability that works today in order to strengthen a guarantee that
      has never applied to it. Recorded as the standing disagreement with the planning gate, and carried into
      the PR description rather than settled quietly.
    3. Stage-only, pending certification — round four's suggestion, and right for safe mode (above) but
      wrong here: certification cannot succeed for these components until (5) exists, so "pending" is
      permanent limbo while deploy_component reports success. A deploy that never takes effect is a worse
      answer than an honest uncertified one.
    4. Certify with appName and mount but omit rootConfigBranchedDatabases — rejected. Loading a
      branch-configured application against the base store is not a rehearsal of the real thing and can
      itself write to base, so it would mint authority a run did not earn. Certifying against the wrong store
      is worse than not certifying.
    5. A validation-scoped branch key that cannot collide with resolveBranchPath(appName, …) — the right
      long-term answer, and what brings these components into the guarantee. New durable-state design; its own
      step.
    6. Accept the durable mutation — rejected: it contradicts the invariant.

    Consequence, stated rather than buried: the guarantee excludes branch-configured components until (4)
    exists — see (5). Nothing they can do today stops working, and nothing they do earns authority — but "no unvalidated
    release is published" does not cover them, and the PR description will say so alongside the restart-revival
    hole. Both narrowings share one shape worth naming: no verdict means no authority, never no verdict means
    no deploy
    .

    Approaches considered

    1. Certification required by prepareApplication(), verdict from a detached one-shot validator —
      proposed. Closes all four call sites; makes .complete mean what recovery already assumes.
    2. Isolated validator wired only into deploy_component's callback — the first draft. Rejected: three
      call sites still mint .complete uncertified.
    3. Delegate to a live serving worker — cheapest, reuses reviewed code, but runs candidate code on a
      thread serving traffic and keeps every hazard the guard exists for.
    4. Child process — the only option that survives native crashes and process.exit, but not
      runtime-equivalent to Harper's in-process database and resource singletons, and a second bootstrap path.
      Held as a successor; the verdict protocol above is reusable unchanged.
    5. Pooled / long-lived validator — amortises spawn cost but carries module and native residue across
      deploys, violating clean-heap-per-verdict.
    6. Drop the gate, validate in-process on main — runs app code on the main thread; a candidate that
      blocks the loop stalls the node.
    7. Do less: activate, watch worker load status, roll back — rejected on the merits: the invalid
      candidate is live in the meantime, and filesystem rollback cannot undo database, network or
      process-global effects.

    Phase-event compatibility

    operations.js is today the only place that emits prepare:done and then the load phase around
    validateComponentLoads. Moving certification inside the preparation would stop those firing, or reorder
    them, for every SSE client that keys progress off phase: load — a wire change with no compatibility
    option weighed in round one. So the phase stream is preserved: certification emits the same load
    start/done phases from wherever it now runs, and prepare:done keeps its position relative to them. Worth
    an explicit regression test, since the operation's error is otherwise the only authority a client has.

    Rollout

    Main-thread deploys gain a load phase, latency, and new failure modes. Replicated deploys can originate on
    a worker, so the nested-validator path needs the parentPort hop rather than a validator spawning a
    validator. Safe mode keeps its no-execution behaviour including preloads. Metrics worth having before this
    is on by default: spawn latency, validation latency, timeout rate, premature exit, cleanup failure.

    Standing disagreement with the planning gate

    Four planning rounds; the first three each produced findings that changed this design materially (the mint
    layer, the branched-database blocker, the deadline basis, the phase stream, the spawn path, safe-mode
    authority). Round four's framing verdict remains better-alternative-exists on one point only: it holds
    that a branch-configured deploy should be refused rather than deployed uncertified, on the grounds that
    compatibility cannot override the invariant the change claims to enforce.

    That is a values judgement, not a factual finding, and it is resolved the other way on purpose: the
    invariant is no verdict means no authority, never no verdict means no deploy. Refusing would delete a
    working capability to widen a guarantee that has never covered it. The gate's objection is recorded here and
    belongs in the PR description under ## For the human reviewer, so a reviewer can overrule it rather than
    discover it.

  5. dawsontoth commented on Sep 8, 2026

    @dawsontoth
    ContributorAuthor

    Step 4 is in review — #2531 (draft)

    Branch claude/harper-2315-step-4-dc4c9b; docs companion HarperFast/documentation#668.

    The premise moved during step 1's review, and the step was re-scoped before implementation. main discards a rejected or failed candidate immediately, and boot recovery removed every owned, journal-less staging directory — so nothing intentionally leaves an unactivated build behind today; only the crash window between .complete and the journal write does. Step 4 therefore lands the bound and the knob as groundwork for step 6, and makes recovery stop destroying the state step 6 will rely on, rather than bounding something that already accumulates.

    What lands:

    • deployment.stagingRetention.maxCount (default 5, 0 keeps none), same coercion shape as payloadRetention.maxSize.
    • A dormant build is .complete + the owner's tree + no journal. Boot recovery keeps dormant builds and bounds them per component (newest by .complete mtime, id tie-break); residue — partial tree, tree already moved live, stale .unsettled — is still removed as before.
    • Removal is decided only under the owner's preparation lock. Boot catalogues unlocked and takes the lock once per over-bound owner, re-deriving only that owner's catalogued directories under it (three review rounds converged here: per-directory locking at boot let healthy components lose the 250 ms probe to sibling threads; a whole-root re-read under the lock recreated the same contention).
    • The deploy path prunes inside the settlement scan it already runs under the lock, before building; drop_component reclaims the dropped component's dormant builds.

    Open decisions carried in the PR's For the human reviewer: count rather than size; .complete mtime as the order; default 5 (retains crash-window residue that used to be swept); a held lock on an over-bound owner defers the component like the residue branch does; and unowned residue (empty directories from a crash during resolve) is deliberately still out of scope because it cannot be told from a build that has not reached extraction.

    For step 6: it must pin the build it is about to activate or prune after staging — the deploy-path prune runs before the build with no notion of a build in use — and it decides whether a stage-only mode wants an explicit dormant marker rather than reusing .complete.

  6. self-assigned this
    on Sep 14, 2026
  7. 16 remaining items

  8. dawsontoth commented on Sep 28, 2026

    @dawsontoth
    ContributorAuthor

    Step 2, re-planned: canary rollout certification

    Replaces the isolated-worker validator shelved on 2026-09-14 (#2476). Proposed by @kriszyp on #2893.

    Why now. #2881 showed that the in-process load check before the swap cannot give a trustworthy verdict. The operations API deploys on the main thread, which never validated, and peers validated inside workers whose intrinsics freeze-after-load had already frozen, so every candidate with a dependency that extends an intrinsic at load (reflect-metadata) was refused on every peer while the origin ran it. #2893 lands first as the 5.3 fix: under the default lockdown no node load-validates a deploy, and none, freeze and ses keep the in-process check. That leaves the default with no validation, which is what this step restores.

    The idea

    Certify a release by rolling it out one worker first:

    1. Keep the previous release restorable (its files and root-config entry), and start one replacement worker on the new release.
    2. The replacement reports whether the component loaded before it is admitted to traffic, and before the old worker it replaces is retired where overlap is supported.
    3. On success, continue the rollout, checking each later replacement too. On failure or timeout, stop the rollout, restore the previous release, and remove or replace any worker running the failed candidate.

    Why this instead of the isolated-worker validator

    The validator was shelved because no host was both serving-equivalent and force-killable. The canary is a real serving worker booted through the normal path, so its load is the boot load (a fresh realm, before the freeze). There is no second bootstrap, no special worker mode, and no separate verdict protocol for a thread that never serves, and it takes no traffic until it reports.

    What it needs that does not exist yet

    • A restorable previous release. That is step 5: re-home the displaced tree as a dormant artifact, so restoring is step 6's activation path pointed at the previous release.
    • Holding a replacement out of traffic until it reports a component-load verdict. Today a worker starts serving once loadRootComponents settles, whether or not each component loaded, so the verdict has to come from component status, not from the worker having started.
    • Platforms without overlap. On Windows, macOS and Bun a replacement only starts after the old worker exits (platformCanPreStartReplacement in server/threads/manageThreads.js), so the canary costs capacity there, and a single-worker node is down for the check.
    • A hung canary under Bun. terminate() still segfaults there, so a canary that blocks its event loop can only be abandoned, not killed. Where overlap exists, the old workers keep serving meanwhile.

    Open questions

    • restart: false deploys start no worker. Canary anyway, or no verdict?
    • Peers: each node canaries independently, so a failure that happens on only one machine still leaves nodes on different versions. Per-node verdicts could come back through the origin the way staging's staged: true markers do, and we need to decide whether one node's failure rolls the others back.
    • The deploy's answer would wait for the canary's verdict, so a restart: true deploy gets slower by one worker boot.
    • What counts as loaded: the component's status reaching healthy, or an explicit readiness signal for components that finish setting up asynchronously.

    Carried from #2893 (merged)

    🤖 Drafted with Claude Code

  9. dawsontoth commented on Oct 2, 2026

    @dawsontoth
    ContributorAuthor

    Step 2, designed: canary rollout certification

    This settles the 2026-09-28 plan and answers its open questions. It went through three cross-model planning rounds (Codex and Gemini). Each round changed the design, and both reviewers' findings from round three are folded in below. It lands as one PR.

    The guarantee, and its limit

    Per node, for controlled starts only. While a release is pending certification:

    • Every worker started that would load it is held, and only one of them is booted at a time: the canary.
    • The canary loads the release through the normal boot path, then waits out of traffic.
    • If the canary rejects the release, or the process dies before the canary decides, no worker is ever started on that release on that node again. The node either restores the previous release or, if none is kept, fails the component closed.

    What it does not cover. Workers already running when the release swaps in keep the code they loaded. Anything they load afterwards still comes from the live directory:

    • static files;
    • a module first imported after the swap;
    • watcher-driven reloads.

    That is how every deploy behaves today. A restore puts the files back, but it cannot unload a module an old worker already imported. #2975 (take the node out of rotation for the whole change) or versioned activation paths would close this gap.

    Decided

    • The origin goes first, for both restart: true and 'rolling'. Studio deploys with 'rolling'.
    • Each node gets its own verdict, and there is no cross-node rollback. A peer that rejects a release the origin certified restores only itself, and the deploy (or rolling job) names it.
    • Later replacements are held and checked too. If one fails, the worker it would have replaced keeps serving and the rollout stops (workersKeptOnOldCode). Nothing is restored.
    • restart: false gets no verdict. This is documented.
    • Pre-swap certification is out of scope. The goal of keeping the node out of rotation while it changes moved to Load-balancer-aware rolling deploys: take each node out of rotation, drain, deploy, rejoin when healthy #2975.

    How it works

    1. The held start. A worker started with a certification instruction boots exactly as any worker does, then holds before listenOnPorts():
      • It reports a verdict built from a private outcome the loader records for that component: executed, skipped, pending or failed. Public health status is not used, because a status update can overwrite a failure, and loadComponent: dev-only or safe mode would report "loaded" without running anything.
      • The verdict also names the deployment id it loaded, read from the tree's provenance marker before and after the load.
      • The worker then waits for main to admit it or shut it down.
    2. The start gate, owned by main.
      • Who is gated: every HTTP worker start that places a pending component, whatever started it (a restart, a crash auto-restart, the isolated reconcile). The gate lives in startWorker, so no caller can bypass it.
      • Canary and waiters: the first such start is the canary. Other starts wait for its decision.
      • Before the swap: while a registration is armed but not yet committed, starts that place the component are deferred. A dead requester is resolved from durable evidence under the preparation lock.
      • Restarts: replacement restarts on main are serialized.
      • Pre-boot: every replacement in a certifying restart pre-boots held while its old worker serves, pool and dedicated alike, on every platform. Admission then follows the platform's handover rule, so a rejection never empties a slot.
      • Deadline: the existing 60-second replacement backstop. A timed-out canary is terminated on Node. On Bun it is abandoned, and that component then refuses certification until the process restarts.
    3. A durable record, .deploy-staging/<id>/.certification.json.
      • Written before the swap, under the preparation lock, after main acknowledges the registration. It names the release this activation will keep as its predecessor.
      • States: pending → certified (the fence holds until the rollout ends) → removed. Or rejected → removed once the restore lands, or kept if there is nothing to restore, in which case the component fails closed on every thread and is never automatically reinstalled.
      • The fence: while a record is pending or certified, every other preparation and drop_component of that component is refused 409. The one exception is an activation of the same id, which joins the in-flight decision.
      • Retention: the predecessor the record names is pinned.
      • At boot, after activation recovery and before installApplications():
        • a pending record left by a dead process is rejected and restored;
        • a certified record is accepted;
        • a rejected record has its restore retried;
        • a record whose release is not live is removed.
      • A failed record write keeps that component's starts blocked, stops the candidate, and still settles every caller.
    4. The restore is step 6's journaled activation of the kept predecessor, guarded so it acts only while the rejected release is still live. Step 5 keeps the rejected tree dormant under its own id, so deployment_id can retry it once its cause is fixed. With no predecessor kept, the component fails closed. That covers a first deploy, deployment_stagingRetention_maxCount: 0, and a tree from a boot install, add_component or before step 5. A rejected first deploy keeps its tree and entry. Re-activating its id re-certifies it.
    5. A deployed tree is not reinstalled over. Today a package deploy's restart re-resolves the package from source, because no deploy writes harper-application-lock.json. That would have the canary certify different bytes from the ones deployed. Instead, installApplications() skips a component whose live tree's provenance names a deployment whose declared entry equals the effective entry in force. Rejection is checked before any automatic install, independently of that check.
    6. The deploy.
      • Origin: prepare, which registers, writes the record and swaps; then wait for the verdict, which is the load phase. On rejection the deploy fails and nothing is replicated. Otherwise it replicates while the rest of the origin's rollout continues.
      • restart: true peers do the same, asking main through the RESTART message. Main replies before it replaces the requesting worker, which it replaces last.
      • 'rolling': the origin certifies, then replicates the release as a stage. A job then activates it on each peer in turn (deploy_component { deployment_id, restart: true }), so a peer's release goes live only at its turn. The job visits every peer and reports each one that did not certify.
      • The response gains certification: certified | uncertified | unavailable | not-requested. A rejection is a failed deploy whose error carries the failures and what was restored.
    7. The in-process load check is retired in every mode. Under none, freeze and ses, a worker-run deploy that throws at load used to be refused. It now activates unless a certifying restart follows. This is documented and covered by a test. The transient-validation apparatus that only that check used is removed in a follow-up.

    Not covered

    • restart: false.
    • Isolation flips.
    • Failures after a component's load completes.
    • Whatever the rejected release did while loading (database writes, files, network, singletons on the canary's worker index).
    • Running workers' later reads of the live directory.
    • Cross-node rollback.
    • Peers on an older build, which restart uncertified. Clusters run one version.

    Verification

    • Unit: real modules and fixture workers, with barriers at each boundary:
      • arming, writing the record, the swap, the decision, and the restore;
      • a worker that crashes during the canary;
      • overlapping restarts and concurrent deploys;
      • skipped loads and safe mode;
      • failed writes, and Bun abandonment.
    • New integration suite:
      • a held canary while the previous release keeps serving;
      • rejection and restore;
      • a rejected first deploy that fails closed and is not reinstalled;
      • the process killed mid-canary;
      • 'rolling' on one node;
      • a package deploy's restart not reinstalling;
      • a lazy import, documenting the stated limit.
    • Cluster, a release gate in a harper-pro companion:
      • an origin rejection replicates nothing;
      • a peer rejection restores only that peer;
      • the rolling job continues past a rejecting peer and an unavailable one;
      • a single-worker peer still answers.

    🤖 Drafted with Claude Code

  10. dawsontoth commented on Oct 6, 2026

    @dawsontoth
    ContributorAuthor

    All seven steps have merged. The last, step 2 (canary rollout certification), landed in #2981 with its docs in HarperFast/documentation#707, and ships in v5.4. The milestone moves to v5.4 to match.

    The two follow-ups still open, #2573 (the 5.2 backport of step 1) and #2659 (activated_from on upgraded installs), now sit under the epic #641, so they stay visible after this closes. The later ideas this sequence left are #2975 (take each node out of rotation for its deploy) and #3025 (server.isPrimaryThread).

  11. modified the milestones: v5.3, v5.4 on Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Fields

Priority

P2

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions