Repository navigation
CRITICAL(deploy): the filtered backend manifest bakes a controller-only constraint path, so provisioning still aborts on every non-manager node #17331
Description
Activity
- addedbugSomething isn't workingSomething isn't working
on Sep 23, 2026 Fix pushed in #17338 (the manifest-path half), head
a4b2e7a0e5.Not closed, and the reason matters more than the fix. The last criterion here is a multi-host provisioning run, which this session cannot perform. The failure was observed on one, the cause is read from the code and from pip's own source, but the fixed run has not happened. Closure needs that run's output, not this merge.
- added 9 commits that reference this issue
on Sep 23, 2026 AC verification against merged
main— 3 of 4 met; AC4 needs a host run I cannot performMerged as
193b182c79. Verified fromorigin/main, quoting the merged file rather than the branch.- The generated manifest's
-c/-rlines resolve on the TARGET.scripts/build-filtered-requirements.sh:53—rewrite_root="${3:-${code_source_dir}}", and thesedrewrites to${rewrite_root}rather than the source root. Omitting the argument keeps both equal, which is whatservices/role_registry.py's post_sync_cmd needs (it reads and installs on the manager). The five delegated generations pass the node's root — e.g.roles/backend/tasks/main.yml:989-992passes{{ code_source_dir }}as the source and{{ autobot.base_dir }}as the rewrite root. - The provisioning path stages what the rewritten path names.
roles/_shared/tasks/stage_shared_manifests.ymlstages both files the script rewrites —dest: {{ autobot.base_dir }}/constraints/shared.txt(:44) anddest: {{ autobot.base_dir }}/requirements.txt(:57) — and nine role task files include it (backendtwice,ai-stack,backend_services,browser,tts-worker, and the threeagent_configtask files), so a role carries its own precondition whatever playbook invoked it. - A guard covers the generated-content case.
repo_tests/ansible_generated_manifest_paths_17331_test.py— four tests: every delegated generation passes a rewrite root, that root is not insidecode_source, the files it names are staged, and each generating role includes the staging task. Floor_MIN_INVOCATIONS_SEENso an empty walk fails rather than passes. Both mutations named in its docstring were executed before merge and both went red. - Verified on a multi-host topology. NOT MET, and not verifiable from merged code. This issue's own argument is that a single-node run cannot reproduce it — the manager is the controller, so the path exists there. The failure was observed on a two-host run; the fixed run has not happened. I cannot produce that evidence from this session, and I am not ticking it on the strength of the guard: the guard proves the rewrite root is passed and staged, not that a fleet node's pip completes.
What would close it: one provisioning run against a non-manager node, with
Install filtered backend requirementsreachingchanged/okinstead ofCould not open constraint file. Paste the recap and this closes.Also unchanged by this merge, and worth stating here so the merge is not read as broader than it is: the 13 husk
dist-infodirectories on the reporting host are untouched.clear_provenance_husksis reachable only from_run_pip_install(the builtin code-sync path);git grep -rn 'venv_reconcile\|clear_provenance_husks' autobot-slm-backend/ansible/returns nothing, and that host is ansible-provisioned. #17339 is that gap.- The generated manifest's
- added a commit that references this issue
on Sep 24, 2026 Closure pass — 3 of 4 criteria met from merged code. Not closed: the host-topology criterion is not measurable now.
Evidence via
git show origin/main:<path>at7adaca8c29.AC1 — the rewrite target is a parameter set to the node-local location — MET, in the second of the two shapes the criterion allowed.
scripts/build-filtered-requirements.sh:53makes the rewrite root a separate argument from the source root:rewrite_root="${3:-${code_source_dir}}"with the contract stated at
:44(<requirements_file> <code_source_dir> [rewrite_root]) and the reason at:27-32. The provisioning path passes the node-local base as that third argument —roles/backend/tasks/main.yml:997-1001:bash {{ code_source_dir }}/scripts/build-filtered-requirements.sh {{ code_source_dir }}/autobot-backend/requirements.txt {{ code_source_dir }} {{ autobot.base_dir }} <- rewrite_root, the path pip opens on the targetAC2 — the provisioning path stages whatever the rewritten path names — MET.
roles/_shared/tasks/stage_shared_manifests.ymlstages both includes onto the node:constraints/shared.txtat:41-47and the rootrequirements.txtat:54, both under{{ autobot.base_dir }}. Included from the backend role at:990-992, guarded by a per-host fact so several roles in one run stage once. The comment at:49-53records why the root manifest had to be staged too —-r ../requirements.txtpulls ~23 root runtime deps, and #11135 filed the silent-drop version of that.AC3 — a guard that covers the generated-content case — MET, and it is a separate check rather than a widened matcher, as the criterion required.
repo_tests/ansible_generated_manifest_paths_17331_test.py:test_every_delegated_generation_names_a_target_side_root(:157),test_the_rewrite_root_is_staged_onto_the_node(:190),test_every_role_that_generates_a_manifest_stages_its_includes(:210), withtest_the_scan_reaches_the_invocations_it_guards(:146) as the floor. Its header at:27documents the blind spot that made a second guard necessary: a guard reading task payloads cannot see a path written into a file at generation time.AC4 — multi-host verification — NOT MEASURABLE NOW. Host mid-update; no reading taken. Unticked deliberately rather than inferred: this is the criterion that #17242's review could not satisfy either, which is how the baked-path half survived.
Held open by AC4 alone. Note for whoever closes: this was carrying no milestone until today, while being
priority: criticaland landed — the state where nobody works it and nobody counts it.- addedblocked: needs-observationNothing but an observation on a running system is left; no session can close itNothing but an observation on a running system is left; no session can close it
on Oct 3, 2026 Labelled
blocked: needs-observation— three of four criteria are ticked and the fourth is not an agent's to satisfyRemaining: "Verified on a multi-host topology — a single-node run cannot reproduce this." The code criteria are all delivered: the generated manifest's
-c/-rlines resolve on the target, the provisioning path stages what the rewritten path names, and a guard covers the generated-content case rather than only the task payload.So this is not open engineering. It is one deploy against more than one node, by someone with the fleet.
Why the label matters here specifically: this is marked CRITICAL, so it reads as urgent unstarted work in every sweep. It is urgent finished work awaiting a confirmation — a very different queue position, and the distinction is invisible without the label.
Grouped with #17243, which is in exactly the same state: four of five ticked, and the fifth is the same multi-host verification. One multi-host deploy answers both.
One chained session discharges these, and what the topology needs to include
Each of #17331, #17243, #17242 already says "one multi-host deploy". The unrecorded part is which hosts and which playbooks, because a deploy that omits one leaves a criterion unmeasured while reading as a pass. The three issues are verified from different playbooks:
Step Run Discharges 1 The documented one-line install on a clean machine, then confirm failed=0#17892 (with #17897 / #17898, as already noted there) 2 Provision a second node through provision-fleet-roleswith at leastbackendandai-stack#17242 AC5, #17331 AC4, #17243 AC5 sites 1-2 3 One monitored Update-All across both nodes #17243 AC5 sites 3-5, #16020 AC6, #12596 What to read back from step 2:
Install filtered backend requirementscompletes on the non-manager node and its first line offiltered-requirements.txtstarts with the node-local base directory, not the checkout path. #17243 records the ai-stack site as a prediction, not an observation (it was never exercised), so the ai-stack role is what turns that into a measurement.What to read back from step 3: the self-update log contains
PLAY [Play 2and aPLAY RECAPlisting the non-manager node. #12596 is open on exactly that: its last comment says the detach fix is merged but no monitored run has shown Play 2 executing. Update-All that skips Play 2 measures nothing for #17243 sites 3-5 or #16020 AC6.Include in the topology if one exists: a node running
npu-worker.roles/npu-worker/tasks/main.ymlstill installs from{{ code_source_dir | default(...) }}(:230,:241,:245) and the #17892 guard records it in_UNRESOLVED_SITES. Whethercode_sourceexists on such a node is the open question in #17334 and in the 2026-10-03 07:15 comment on #17892. Not determined: whether the fleet has one.Also useful, not required: a node with
slm-agent+vnc(#16020 AC6) and one withredisonly (#16020 AC2). I could not determine from the repo whether the failed non-manager node in the 2026-09-23 run was the same machine as the vnc node.Host observation — the generated manifest on this node no longer bakes a controller-only path. It bakes the deploy base directory, and the file it names is present there.
Read-only, from a live deployment whose checkout is at current
main:autobot-backend/filtered-requirements.txton disk begins with
-c <base_dir>/constraints/shared.txt # single source of truth for shared versions (#10524)
— the base directory, not<base_dir>/code_source/constraints/. The path this issue reported baked into the artifact is gone from the artifact.<base_dir>/constraints/shared.txtexists on this node and is byte-identical to the checkout's copy.<base_dir>/requirements.txtis present too.- The mechanism is visible in the deployed role:
roles/backend/tasks/main.ymlincludes_shared/tasks/stage_shared_manifests.ymlbefore the generation step, and passes{{ autobot.base_dir }}as the script's fourth argument (therewrite_rootthe deployedscripts/build-filtered-requirements.shgrew for this issue). The same pair appears inroles/ai-stack/tasks/main.yml. - The most recent builtin-updater run on this node completed
failed=0for both the manager and the controller play.
What it discharges: the artifact-level claim. The rewritten
-cline no longer names the controller's checkout, and the staging task that makes the new target real is in place and demonstrably effective on this node.What it does not discharge: the issue's actual failure case. This is a single-node deployment, so the node running the task is the controller — exactly the configuration the issue says cannot expose the defect. A second node is registered in the fleet, but it carries only the
vncandslm-agentroles, so it never reaches the backend or ai-stack pip steps. I could not determine whether a non-manager backend node now provisions cleanly; that still needs a run against one.
#17242moved the generation of the filtered backend manifest to the controller. It did notmove the path that generation writes into the file, and pip reads that file on the target.
Provisioning still aborts on every non-manager node, at the next task instead of the same one.
Evidence — a run from this evening, not inferred
PLAY RECAP for that run: the non-manager node
failed=1, and the managerfailed=1for anunrelated reason (filed separately).
The installed tree already carries
#17242's fix —roles/backend/tasks/main.ymlcontains theCreate filtered backend requirements (on the controller, #17242)task withdelegate_to: localhost, and the deployed checkout is at currentmain. So this is not#17242recurring: it is the layer#17242did not reach.Root cause
scripts/build-filtered-requirements.sh:45rewrites the sibling-relative include to thecontroller's checkout path:
The generated file therefore begins:
That line is then copied verbatim onto the target by
roles/backend/tasks/main.yml(Write filtered backend requirements to the target (#17242))and consumed by the
pip:task below it, which runs on the target.code_sourceis thecontroller's git checkout; nothing syncs it to a node — the invariant
#17242established.So the class is unchanged: a controller-only path reaching a target.
#17242fixed the executorand left the payload.
The header of
scripts/build-filtered-requirements.sh:43already documents this exact string asa provisioning abort, and
#17243already documented cases 3–5 failing this way. The fix shapeis also already in the tree:
playbooks/update-all-nodes.ymlstagesconstraints/shared.txtto
<base_dir>/constraints/shared.txtand constrains against the local copy. The provisioningpath (
provision-fleet-roles) does no such staging.Why the guard did not catch it
repo_tests/ansible_code_source_delegation_17243_test.pyscans a task's payload text forcode_source. The pip task's payload isrequirements: {{ backend_code_dir }}/filtered-requirements.txt— no
code_sourcein it. The controller-only path is inside the file's contents, produced atrun time by a shell script. A guard that reads task payloads cannot see a path that a generator
writes.
Acceptance criteria
-c(and-r) lines resolve on the target, not on thecontroller — either by staging
constraints/shared.txt(and the rootrequirements.txt)to a node-local path and rewriting to that, or by making the script's rewrite target a
parameter the provisioning path sets to the node-local location.
already does for its own pip steps.
that is not staged fails a test. The existing delegation guard is documented as unable to
see this, so a new check is needed rather than a widened matcher.
it survived
#17242's review.Related
#17242— the executor half of this defect, fixed and merged#17243— the sweep; its cases 3–5 are the same "Could not open constraint file" shape#17246— seven pip sites reading manifests out ofcode_source, still baselined#17317— the log reader truncated the first copy of this error mid-word, which is why thecause was unreadable when it was first hit