Skip to content

CRITICAL(deploy): the filtered backend manifest bakes a controller-only constraint path, so provisioning still aborts on every non-manager node #17331

Description

@mrveiss

#17242 moved the generation of the filtered backend manifest to the controller. It did not
move the path that generation writes into the file, and pip reads that file on the target.
Provisioning still aborts on every non-manager node, at the next task instead of the same one.

Evidence — a run from this evening, not inferred

TASK [backend : Install filtered backend requirements]
fatal: [<non-manager-node>]: FAILED! => {"changed": false,
  "cmd": [".../autobot-backend/venv/bin/pip3", "install", "-r",
          ".../autobot-backend/filtered-requirements.txt"],
  "msg": "\n:stderr: ERROR: Could not open constraint file: [Errno 2] No such file or
          directory: '<base_dir>/code_source/constraints/shared.txt'\n"}

PLAY RECAP for that run: the non-manager node failed=1, and the manager failed=1 for an
unrelated reason (filed separately).

The installed tree already carries #17242's fix — roles/backend/tasks/main.yml contains the
Create filtered backend requirements (on the controller, #17242) task with
delegate_to: localhost, and the deployed checkout is at current main. So this is not
#17242 recurring: it is the layer #17242 did not reach.

Root cause

scripts/build-filtered-requirements.sh:45 rewrites the sibling-relative include to the
controller's checkout path:

sed -E "s|^-c (\.\./)+constraints/|-c ${code_source_dir}/constraints/|; \
        s|^-r (\.\./)+requirements\.txt|-r ${code_source_dir}/requirements.txt|"

The generated file therefore begins:

-c <base_dir>/code_source/constraints/shared.txt  # single source of truth for shared versions (#10524)

That line is then copied verbatim onto the target by
roles/backend/tasks/main.yml (Write filtered backend requirements to the target (#17242))
and consumed by the pip: task below it, which runs on the target. code_source is the
controller's git checkout; nothing syncs it to a node — the invariant #17242 established.

So the class is unchanged: a controller-only path reaching a target. #17242 fixed the executor
and left the payload.

The header of scripts/build-filtered-requirements.sh:43 already documents this exact string as
a provisioning abort, and #17243 already documented cases 3–5 failing this way. The fix shape
is also already in the tree: playbooks/update-all-nodes.yml stages constraints/shared.txt
to <base_dir>/constraints/shared.txt and constrains against the local copy. The provisioning
path (provision-fleet-roles) does no such staging.

Why the guard did not catch it

repo_tests/ansible_code_source_delegation_17243_test.py scans a task's payload text for
code_source. The pip task's payload is requirements: {{ backend_code_dir }}/filtered-requirements.txt
— no code_source in it. The controller-only path is inside the file's contents, produced at
run time by a shell script. A guard that reads task payloads cannot see a path that a generator
writes.

Acceptance criteria

  • The generated manifest's -c (and -r) lines resolve on the target, not on the
    controller — either by staging constraints/shared.txt (and the root requirements.txt)
    to a node-local path and rewriting to that, or by making the script's rewrite target a
    parameter the provisioning path sets to the node-local location.
  • The provisioning path stages whatever the rewritten path names, as the updater playbook
    already does for its own pip steps.
  • A guard covers the generated-content case, not only the task-payload case: a rewrite target
    that is not staged fails a test. The existing delegation guard is documented as unable to
    see this, so a new check is needed rather than a widened matcher.
  • Verified on a multi-host topology — a single-node run cannot reproduce this, which is why
    it survived #17242's review.

Related

  • #17242 — the executor half of this defect, fixed and merged
  • #17243 — the sweep; its cases 3–5 are the same "Could not open constraint file" shape
  • #17246 — seven pip sites reading manifests out of code_source, still baselined
  • #17317 — the log reader truncated the first copy of this error mid-word, which is why the
    cause was unreadable when it was first hit

Activity

  1. mrveiss commented on Sep 23, 2026

    @mrveiss
    OwnerAuthor

    Fix pushed in #17338 (the manifest-path half), head a4b2e7a0e5.

    Not closed, and the reason matters more than the fix. The last criterion here is a multi-host provisioning run, which this session cannot perform. The failure was observed on one, the cause is read from the code and from pip's own source, but the fixed run has not happened. Closure needs that run's output, not this merge.

  2. mrveiss commented on Sep 24, 2026

    @mrveiss
    OwnerAuthor

    AC verification against merged main — 3 of 4 met; AC4 needs a host run I cannot perform

    Merged as 193b182c79. Verified from origin/main, quoting the merged file rather than the branch.

    • The generated manifest's -c/-r lines resolve on the TARGET. scripts/build-filtered-requirements.sh:53 — rewrite_root="${3:-${code_source_dir}}", and the sed rewrites to ${rewrite_root} rather than the source root. Omitting the argument keeps both equal, which is what services/role_registry.py's post_sync_cmd needs (it reads and installs on the manager). The five delegated generations pass the node's root — e.g. roles/backend/tasks/main.yml:989-992 passes {{ code_source_dir }} as the source and {{ autobot.base_dir }} as the rewrite root.
    • The provisioning path stages what the rewritten path names. roles/_shared/tasks/stage_shared_manifests.yml stages both files the script rewrites — dest: {{ autobot.base_dir }}/constraints/shared.txt (:44) and dest: {{ autobot.base_dir }}/requirements.txt (:57) — and nine role task files include it (backend twice, ai-stack, backend_services, browser, tts-worker, and the three agent_config task files), so a role carries its own precondition whatever playbook invoked it.
    • A guard covers the generated-content case. repo_tests/ansible_generated_manifest_paths_17331_test.py — four tests: every delegated generation passes a rewrite root, that root is not inside code_source, the files it names are staged, and each generating role includes the staging task. Floor _MIN_INVOCATIONS_SEEN so an empty walk fails rather than passes. Both mutations named in its docstring were executed before merge and both went red.
    • Verified on a multi-host topology. NOT MET, and not verifiable from merged code. This issue's own argument is that a single-node run cannot reproduce it — the manager is the controller, so the path exists there. The failure was observed on a two-host run; the fixed run has not happened. I cannot produce that evidence from this session, and I am not ticking it on the strength of the guard: the guard proves the rewrite root is passed and staged, not that a fleet node's pip completes.

    What would close it: one provisioning run against a non-manager node, with Install filtered backend requirements reaching changed/ok instead of Could not open constraint file. Paste the recap and this closes.

    Also unchanged by this merge, and worth stating here so the merge is not read as broader than it is: the 13 husk dist-info directories on the reporting host are untouched. clear_provenance_husks is reachable only from _run_pip_install (the builtin code-sync path); git grep -rn 'venv_reconcile\|clear_provenance_husks' autobot-slm-backend/ansible/ returns nothing, and that host is ansible-provisioned. #17339 is that gap.

  3. added this to the v0.9.0 milestone on Sep 28, 2026
  4. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    Closure pass — 3 of 4 criteria met from merged code. Not closed: the host-topology criterion is not measurable now.

    Evidence via git show origin/main:<path> at 7adaca8c29.

    AC1 — the rewrite target is a parameter set to the node-local location — MET, in the second of the two shapes the criterion allowed. scripts/build-filtered-requirements.sh:53 makes the rewrite root a separate argument from the source root:

    rewrite_root="${3:-${code_source_dir}}"

    with the contract stated at :44 (<requirements_file> <code_source_dir> [rewrite_root]) and the reason at :27-32. The provisioning path passes the node-local base as that third argument — roles/backend/tasks/main.yml:997-1001:

    bash {{ code_source_dir }}/scripts/build-filtered-requirements.sh
         {{ code_source_dir }}/autobot-backend/requirements.txt
         {{ code_source_dir }}
         {{ autobot.base_dir }}          <- rewrite_root, the path pip opens on the target
    

    AC2 — the provisioning path stages whatever the rewritten path names — MET. roles/_shared/tasks/stage_shared_manifests.yml stages both includes onto the node: constraints/shared.txt at :41-47 and the root requirements.txt at :54, both under {{ autobot.base_dir }}. Included from the backend role at :990-992, guarded by a per-host fact so several roles in one run stage once. The comment at :49-53 records why the root manifest had to be staged too — -r ../requirements.txt pulls ~23 root runtime deps, and #11135 filed the silent-drop version of that.

    AC3 — a guard that covers the generated-content case — MET, and it is a separate check rather than a widened matcher, as the criterion required. repo_tests/ansible_generated_manifest_paths_17331_test.py: test_every_delegated_generation_names_a_target_side_root (:157), test_the_rewrite_root_is_staged_onto_the_node (:190), test_every_role_that_generates_a_manifest_stages_its_includes (:210), with test_the_scan_reaches_the_invocations_it_guards (:146) as the floor. Its header at :27 documents the blind spot that made a second guard necessary: a guard reading task payloads cannot see a path written into a file at generation time.

    AC4 — multi-host verification — NOT MEASURABLE NOW. Host mid-update; no reading taken. Unticked deliberately rather than inferred: this is the criterion that #17242's review could not satisfy either, which is how the baked-path half survived.

    Held open by AC4 alone. Note for whoever closes: this was carrying no milestone until today, while being priority: critical and landed — the state where nobody works it and nobody counts it.

  5. modified the milestones: v0.9.0, v0.9-host on Sep 28, 2026
  6. added
    blocked: needs-observationNothing but an observation on a running system is left; no session can close it
    on Oct 3, 2026
  7. mrveiss commented on Oct 3, 2026

    @mrveiss
    OwnerAuthor

    Labelled blocked: needs-observation — three of four criteria are ticked and the fourth is not an agent's to satisfy

    Remaining: "Verified on a multi-host topology — a single-node run cannot reproduce this." The code criteria are all delivered: the generated manifest's -c/-r lines resolve on the target, the provisioning path stages what the rewritten path names, and a guard covers the generated-content case rather than only the task payload.

    So this is not open engineering. It is one deploy against more than one node, by someone with the fleet.

    Why the label matters here specifically: this is marked CRITICAL, so it reads as urgent unstarted work in every sweep. It is urgent finished work awaiting a confirmation — a very different queue position, and the distinction is invisible without the label.

    Grouped with #17243, which is in exactly the same state: four of five ticked, and the fifth is the same multi-host verification. One multi-host deploy answers both.

  8. mrveiss commented on Oct 3, 2026

    @mrveiss
    OwnerAuthor

    One chained session discharges these, and what the topology needs to include

    Each of #17331, #17243, #17242 already says "one multi-host deploy". The unrecorded part is which hosts and which playbooks, because a deploy that omits one leaves a criterion unmeasured while reading as a pass. The three issues are verified from different playbooks:

    Step Run Discharges
    1 The documented one-line install on a clean machine, then confirm failed=0 #17892 (with #17897 / #17898, as already noted there)
    2 Provision a second node through provision-fleet-roles with at least backend and ai-stack #17242 AC5, #17331 AC4, #17243 AC5 sites 1-2
    3 One monitored Update-All across both nodes #17243 AC5 sites 3-5, #16020 AC6, #12596

    What to read back from step 2: Install filtered backend requirements completes on the non-manager node and its first line of filtered-requirements.txt starts with the node-local base directory, not the checkout path. #17243 records the ai-stack site as a prediction, not an observation (it was never exercised), so the ai-stack role is what turns that into a measurement.

    What to read back from step 3: the self-update log contains PLAY [Play 2 and a PLAY RECAP listing the non-manager node. #12596 is open on exactly that: its last comment says the detach fix is merged but no monitored run has shown Play 2 executing. Update-All that skips Play 2 measures nothing for #17243 sites 3-5 or #16020 AC6.

    Include in the topology if one exists: a node running npu-worker. roles/npu-worker/tasks/main.yml still installs from {{ code_source_dir | default(...) }} (:230, :241, :245) and the #17892 guard records it in _UNRESOLVED_SITES. Whether code_source exists on such a node is the open question in #17334 and in the 2026-10-03 07:15 comment on #17892. Not determined: whether the fleet has one.

    Also useful, not required: a node with slm-agent + vnc (#16020 AC6) and one with redis only (#16020 AC2). I could not determine from the repo whether the failed non-manager node in the 2026-09-23 run was the same machine as the vnc node.

  9. mrveiss commented on Oct 3, 2026

    @mrveiss
    OwnerAuthor

    Host observation — the generated manifest on this node no longer bakes a controller-only path. It bakes the deploy base directory, and the file it names is present there.

    Read-only, from a live deployment whose checkout is at current main:

    • autobot-backend/filtered-requirements.txt on disk begins with
      -c <base_dir>/constraints/shared.txt # single source of truth for shared versions (#10524)
      — the base directory, not <base_dir>/code_source/constraints/. The path this issue reported baked into the artifact is gone from the artifact.
    • <base_dir>/constraints/shared.txt exists on this node and is byte-identical to the checkout's copy. <base_dir>/requirements.txt is present too.
    • The mechanism is visible in the deployed role: roles/backend/tasks/main.yml includes _shared/tasks/stage_shared_manifests.yml before the generation step, and passes {{ autobot.base_dir }} as the script's fourth argument (the rewrite_root the deployed scripts/build-filtered-requirements.sh grew for this issue). The same pair appears in roles/ai-stack/tasks/main.yml.
    • The most recent builtin-updater run on this node completed failed=0 for both the manager and the controller play.

    What it discharges: the artifact-level claim. The rewritten -c line no longer names the controller's checkout, and the staging task that makes the new target real is in place and demonstrably effective on this node.

    What it does not discharge: the issue's actual failure case. This is a single-node deployment, so the node running the task is the controller — exactly the configuration the issue says cannot expose the defect. A second node is registered in the fleet, but it carries only the vnc and slm-agent roles, so it never reaches the backend or ai-stack pip steps. I could not determine whether a non-manager backend node now provisions cleanly; that still needs a run against one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions