Repository navigation
bug(code-sync): update-all reports 'completed' while co-located component deploys failed — the self-node still reads up_to_date #16640
Description
Activity
- addedbugSomething isn't workingSomething isn't working
on Sep 13, 2026 The browser-service failure above is filed as 16641 (venv creation needs a
virtualenvexecutable). The npu-worker failure is #15733.Another occurrence on 2026-09-14, in an update-all to 7b49b52. The run ended with
Fleet stage complete: Updated 2/2 nodes, and both nodes reportup_to_date. During the same run, two co-located component deploys on the SLM manager failed:browser: pip failed because thevirtualenvexecutable is missing (fix(ansible): 8 pip tasks create venvs via the 'virtualenv' executable — browser-service deploy fails on hosts that have only python3 -m venv #16641).npu-worker: "Install worker dependencies from its manifest" failed with../autobot_shared is not a valid editable requirement.
Neither failure appears in the update-all result, so the operator sees success while two services run on dependencies from before the update.
#16789 merged (
e49d3d077) — this issue stays open, and was never referenced by itPR #16789 landed on
mainase49d3d077. Its branch was namedissue-16717-16640-16310-code-sync-batch,
so this issue was in its intended scope, but the PR body referenced onlyCloses #16717and
Refs #16310— this issue appears nowhere in it. So it neither auto-closed nor got a tracking link,
and would have gone quiet had the branch name not named it.Recording that rather than silently closing it, for two reasons.
I have not verified the acceptance criteria. Merging is not closing, and none of the four criteria
here have been checked against the merged code. Two of them — the node's reported status not reading
up_to_datewhile a component failed, and the maintenance GUI showing the partial result rather than a
green "completed" — are behavioural claims about a running system, not things a diff can settle.What the next person needs to do, in order:
- Check whether
e49d3d077actually contains work for this issue at all, or whether the branch name
over-claimed its scope. The batch delivered the bug(code-sync): a forced drift resync of autobot-slm-frontend deletes the live bundle and its rollback (dist-*, current, previous are not excluded) #16717 drift-resync fix; whether it also covers
partial-failure reporting is unestablished. - If it does, tick each criterion with the
file:symbolevidence. - If it does not, this issue is simply still open and the branch name was aspirational.
Either answer is fine; leaving it ambiguous is not.
Cross-reference: this is the batch-closure failure mode where a PR carrying several issues closes only
the ones named in its body. The others land silently with their issues open, and nothing surfaces it
until someone audits. Worth noting on any future batch: the branch name is not a closing keyword.- Check whether
Closure: 4 of 4 verified against
origin/main(1edec3064c)Landed via the code-sync batch #16789 (with #16717 / #16310). The co-located procedure moved to
api/_colocated_role_procedures.pyto stay undercode_sync.py's size ceiling.Criterion Evidence A failed co-located deploy ends its stage and the update-all job non-success, naming each component and its reason autobot-slm-backend/api/code_sync.py:5086-5091: onfailed_colocated,stage.status = _StageStatus.PARTIALwithstage.message = "SLM already at <sha> — co-located failed: <name>: <reason>; …"(the"<name>: <reason>"lines come from_colocated_role_procedures.py:36).:5440-5444:if slm_stage.status == _StageStatus.PARTIAL: job.status = "partial".PARTIAL = "partial"at:4588.The node's code status does not read up_to_datewhile a component failedthe test pins it: advance.assert_not_awaited(), "code_status must not advance to up_to_date over a failed component"(tests/api/test_code_sync_update_all_colocated_failure_16640.py:118)A test drives a co-located deploy to failure and asserts stage/job status and per-component detail the same file: test_a_failed_colocated_component_leaves_the_stage_partial_with_the_reason(:109),test_already_current_dispatch_reports_partial_not_completed(:139), plus contrasts:121,:151,:162. 5 passed (locally, on the deployed SLM interpreter, tree identical to main for these files).The maintenance GUI shows the partial result with component and reason, not a green "completed" autobot-slm-frontend/src/views/CodeSyncView.vue:partialgets its own amber style and label (:174,:193);colocatedPartial(:198-205) separates this cause from fleet skips, with its own banner (:1114); each stage'sstage.messageis rendered (:1035-1039), and that is where the component and reason live. Read from source, not seen rendered.Closing.
Host evidence, 2026-09-13 (Riga time), local two-node install:
POST /code-sync/update-allfinished with statuscompleted, andslm_self_updatereadcurrent("SLM already at b688b3e — no restart needed"). But in that same stage, two co-located component deploys failed:The ansible recaps for those two runs read
failed=1:Failed to find required executable "virtualenv".../autobot_sharedas an editable requirement (tech-debt(deploy): three manifests use relative editables that only resolve from an implicit cwd #15733).The node list still reports the self-node
code_status: up_to_date, so nothing tells the operator that two components are still on their previous code. The failures are visible only in the server log.Acceptance criteria
partial), naming each failed component and the reason line from its output.up_to_datewhile any of its components failed to deploy at the target revision.