Part of #641.
Problem
A certified rolling deploy (5.4.0, #2981) happens in two steps.
- The node that receives the deploy stages the release on its peers, certifies it and activates it locally, then responds. The deployment record it wrote is finished at this point, with status
success.
- A
restart_service job then activates the release on each peer in turn (activateDeploymentOnPeers in bin/restart.ts). Each peer's activation runs with _deploymentId, so it writes no record of its own (components/operations.js). Each node's outcome (activated[]: node, certification, message) lands only in the job's result.
So get_deployment and list_deployments report success for a release that a node rejected, or never reached. They keep reporting it after the job ends, because the job isn't the history anyone reads. The documentation tells CI to "wait on the job, not the record", yet it still recommends list_deployments for "what exactly is deployed right now".
Proposal
- When the job finishes on a peer, write that node's outcome into the record's
peer_results entry: whether it activated, its certification, and its error if any. Keep the staging outcome the entry already holds.
- While the job runs, give the record a status that says the rollout is still in progress (for example
rolling_out). Finish it as success only when every node took the release. A rollout that some node didn't take ends with a distinct status (for example partial) that names the nodes.
- Record the job's id on the record, so a client holding a
deployment_id can find the job.
The job runs on the node that wrote the record.
Done when
On a two-node 5.4 cluster where the second node's canary rejects the release, the record shows:
- the first node certified and activated;
- the second node rejected, with its reason;
- an overall status other than
success.
Part of #641.
Problem
A certified rolling deploy (5.4.0, #2981) happens in two steps.
success.restart_servicejob then activates the release on each peer in turn (activateDeploymentOnPeersinbin/restart.ts). Each peer's activation runs with_deploymentId, so it writes no record of its own (components/operations.js). Each node's outcome (activated[]: node, certification, message) lands only in the job's result.So
get_deploymentandlist_deploymentsreportsuccessfor a release that a node rejected, or never reached. They keep reporting it after the job ends, because the job isn't the history anyone reads. The documentation tells CI to "wait on the job, not the record", yet it still recommendslist_deploymentsfor "what exactly is deployed right now".Proposal
peer_resultsentry: whether it activated, itscertification, and its error if any. Keep the staging outcome the entry already holds.rolling_out). Finish it assuccessonly when every node took the release. A rollout that some node didn't take ends with a distinct status (for examplepartial) that names the nodes.deployment_idcan find the job.The job runs on the node that wrote the record.
Done when
On a two-node 5.4 cluster where the second node's canary rejects the release, the record shows:
success.