Skip to content

nginx cannot start on a node with no TLS cert: five roles consume the shared keypair, four provision it #16020

Description

@mrveiss

Symptom

nginx is failing to start on the VNC node and has been since at least 2026-09-07 23:18 — six consecutive failed restarts, each blocking the manage-service playbook.

nginx[1938265]: [emerg] cannot load certificate "/etc/autobot/certs/server-cert.pem":
                BIO_new_file() failed (SSL: error:80000002:system library::No such file or directory)
nginx: configuration file /etc/nginx/nginx.conf test failed
systemd[1]: nginx.service: Control process exited, code=exited, status=1/FAILURE
TASK [SVC | Re-raise the original failure (#15879)]
fatal: [node]: FAILED! => "nginx could not be restarted"

The nginx config references a certificate the node does not have. ExecStartPre=nginx -t fails, so the unit can never come up, and every restart attempt re-fails identically.

The defect is a population, not an instance

Counted across autobot-slm-backend/ansible/roles/:

roles
consume /etc/autobot/certs/server-cert.pem backend, frontend, redis, slm_manager, vnc
provision it (openssl req) backend, frontend, slm_manager, vnc

redis consumes the shared keypair and never provisions it. It is the same latent failure as the nginx one, waiting for a node whose role set does not happen to include a provisioner.

And the nginx role itself does exactly three things — refresh apt, install nginx, start the service. It starts a daemon whose config may reference a cert it never checks for.

Why it survived

Four roles each carry their own copy of the same openssl req block. The vnc role's own comment records the pattern and its reason:

Shared node keypair, also provisioned by roles/backend and setup-internal-ca. Kept here so a vnc-only node still ends up with one.

So the invariant "any node reaching a cert consumer has a cert" is maintained by each consumer separately remembering to, with nothing enforcing it. It holds for as long as nobody adds a consumer without a copy. Someone added redis.

A duplicated guarantee is one that fails silently the first time it is not duplicated, and it fails on whichever node has the unlucky role set — which is why this shows up on the VNC node and not on the control node.

Fix direction

One shared task file — roles/_shared/tasks/ensure_node_tls_cert.yml, alongside the six that already live there — included by every consumer, replacing four hand-copied blocks with one. nginx includes it before starting the service; redis gets it for the first time.

Acceptance criteria

  • The cert-ensure block exists once and every consuming role includes it; no role carries a private copy
  • redis provisions the cert it consumes — verified on a node whose roles are redis-only
  • The nginx role does not start nginx before the certs its config references exist
  • A test asserts the two sets are equal: every role that references /etc/autobot/certs/ includes the shared task. A guard over the roles tree, so a sixth consumer cannot be added without one
  • The failure mode is diagnosable: if the cert is genuinely absent at start time, the error names the missing path and the role expected to provide it, rather than nginx -t exit 1
  • The update path makes sure the shared keypair exists on every node running a service that references /etc/autobot/certs/. Today no role that update-all runs on this node does so. Verified by an update-all run on a slm-agent+vnc node, whose nginx starts, or is correctly absent.

The fourth criterion is what stops this recurring. The current invariant lives in a comment; an invariant a sweep can read is a control, and prose is not.

Blocks the operator's VNC node. Related: #16022 (setup cannot get past provision) may be this same failure surfacing in the provisioning path — filed separately because the log lines proving the nginx failure all come from the service-management playbook, not a provisioning one. Also #16019, #16021.

Activity

  1. mrveiss commented on Sep 8, 2026

    @mrveiss
    OwnerAuthor

    Closed on merged base, per criterion. Every quote is from git show origin/Dev_new_gui:<path> after #16023 merged — not from the PR diff.

    AC1 — the cert-ensure block exists once and every consuming role includes it. ansible/_shared/tasks/ensure_node_tls_cert.yml is present on base. Consumers and provisioners now match exactly:

    roles
    read /etc/autobot/certs/ backend, frontend, nginx, redis, slm_manager, vnc
    ensure it exists the same six

    Gap before: ['redis']. Gap after: [].

    AC2 — redis provisions the cert it consumes. roles/redis/tasks/main.yml includes the shared task, before installing Redis Stack.

    AC3 — nginx does not start before the certs exist. roles/nginx/tasks/main.yml on base: the include is at line 58, state: started at line 65. Ordering, not merely presence — presence alone would have passed a version that ensured the cert after starting the service, which is the bug.

    AC4 — a test asserts the two sets are equal. repo_tests/tls_cert_consumers_provision_it_16020_test.py compares them as sets, with a floor on roles examined so a collapsed sweep cannot report "no gap" the way a fixed tree does. Non-vacuity is not theoretical: the guard fails on the tree that shipped.

    AC5 — the failure is diagnosable — NOT met, and I am not ticking it. The shared task names the missing path and the roles that provide it in its header, so a reader who reaches the file learns why. But a node that somehow still lacks the cert fails at nginx -t with exit 1 exactly as before; nothing added a preflight that reports the missing path at the moment of failure. That criterion is unmet and stays unmet.

    What this does not fix

    The keypair now exists on every node that reaches a consumer. A node already in the broken state is not repaired until Ansible next runs against it — this changes what provisioning produces, not what a failed node currently has. The VNC node needs a provisioning run to pick it up.

    What it taught

    Four roles carried hand-copied openssl blocks and the invariant was held by each consumer separately remembering to — the vnc role's comment says so outright: "Kept here so a vnc-only node still ends up with one." A duplicated guarantee fails silently the first time it is not duplicated, on whichever node has the unlucky role set.

    And I nearly wrote a fifth copy. _shared/tasks/generate_self_signed_cert.yml already existed from #12181, extracted from rotate-certs.yml, with exactly one consumer. The new task delegates to it. An extraction with one consumer looks like dead weight and is actually an unadopted standard — the same shape as pre-commit-no-direct-redis using tokenize while five other guards grep (#16011).

    Merged as #16023. Remaining live-deployment issues: #16019 (#16027 open), #16021, #16022.

  2. mrveiss commented on Sep 8, 2026

    @mrveiss
    OwnerAuthor

    Reopening the finding: the fix is correct and does not reach the failure path

    Still failing on the live deployment, hours after #16023 merged and deployed. Same node, same cert, three more times today:

    Sep 08 12:36:35 b9a29e04 nginx[1957982]: [emerg] cannot load certificate "/etc/autobot/certs/server-cert.pem"
    Sep 08 16:39:41 b9a29e04 nginx[1976462]: [emerg] cannot load certificate ...
    Sep 08 16:40:49 b9a29e04 nginx[1978159]: [emerg] cannot load certificate ...
    

    The deployment is current — .deployed_commit is cbe0e646c841, identical to base — so this is not a sync lag. The fix is deployed and the failure continues.

    Why

    ensure_node_tls_cert.yml is included by the nginx and redis roles. The failing path runs manage-service.yml, which has zero roles: — it is a bare service-state play invoked by api/orchestration.py, services/reconciler.py and services/playbook_executor.py.

    So the include never fires on the path that actually restarts nginx. I fixed provisioning; the operator's failure is in service management, and those are different playbooks.

    I said at closure that "a node already in the broken state is not repaired until Ansible next runs against it" — that was right about the mechanism and wrong about the consequence. Ansible has run against it, repeatedly, via a play that cannot repair it. Waiting does not fix this node; nothing in the current design will.

    Why the acceptance criteria did not catch it

    Criterion 3 was "the nginx role does not start nginx before the certs its config references exist" — verified, line 58 before line 65. True, and scoped to the role. The criteria never asked "can the cert be missing on a path that starts nginx?", and manage-service starts nginx while running no roles at all.

    That is this repository's recurring shape once more: a guard that measures the dimension the bug left. The role-level assertion holds; the defect moved to the playbook that bypasses roles.

    What the fix has to be

    Not a wider role include — the restart path runs no roles by construction. Either:

    1. manage-service.yml gains a pre-task that ensures the shared keypair when the target service is one that consumes it, or
    2. the cert becomes a node-level invariant checked wherever a service is started, not a role-level one.

    (1) is narrower and honest about what it covers; (2) is the version that stops the next consumer being missed, and matches how this issue was framed — five consumers, four provisioners.

    Immediate state

    This node is down and will stay down until something provisions it. The three failing restarts today were the reconciler or an operator retrying, each re-failing identically.

    Reopening rather than filing a follow-up: the issue's title is exactly what is still true — nginx cannot start on a node with no TLS cert — and closing it on a role-scoped fix is what let the live failure continue unnoticed for nine hours.

  3. reopened this on Sep 8, 2026
  4. mrveiss commented on Sep 8, 2026

    @mrveiss
    OwnerAuthor

    Live evidence, read-only, 2026-09-08

    The failure is still current and the cause is now recorded rather than inferred. From the SLM store, services.extra_data for nginx on the VNC node, last_checked 17:47 UTC today:

    [emerg] cannot load certificate "/etc/autobot/certs/server-cert.pem":
    BIO_new_file() failed (SSL: ... No such file or directory: calling fopen(...))
    

    nginx -t runs in ExecStartPre, so this is not a degraded start — the unit cannot come up.

    Auto-remediation has given up six times in two days, each with the same terminal event:

    Service nginx on VNC requires human intervention after 3 failed restart attempts
      2026-09-07 20:26 · 09-08 05:44 · 09:38 · 13:47 · 16:08 · 17:45 (UTC)
    

    Fleet-wide the status picture is otherwise clean: stopped 156 · running 77 · completed 50 · failed 4 · crash-loop 1, and zero unknown — which is separate host confirmation that the #16019/#16027 collector fix is live and working (completed 50 is the active/exited mapping doing its job).

    Why the earlier close was wrong, precisely

    _shared/tasks/ensure_node_tls_cert.yml is included by exactly two roles:

    roles/nginx/tasks/main.yml:58
    roles/redis/tasks/main.yml:67
    

    nginx on this node is not owned by roles/nginx. It is owned by roles/vnc — and roles/vnc/tasks/main.yml still carries its own hand-rolled openssl block at :246, one of the four duplicates the shared task's docstring was written to consolidate. It references server-cert.pem (:242-254) and includes the shared task nowhere.

    So the fix landed in the roles that were already correct and missed the one that was failing. The consumers that reference the cert path without including the shared task are roles/vnc, roles/backend, roles/frontend, and roles/slm_manager.

    The second gap, which is why nobody could repair it by hand

    The operator-facing repair path is manage-service.yml, which has tasks: and no roles: — so it can restart nginx but can never establish the cert nginx needs. The rescue block added in #15879 works exactly as designed here: the [emerg] line above reached the controller log instead of staying on the node. It made the failure legible; it cannot make it repairable.

    And the two paths that could re-run provisioning both refuse this node:

    path gate result
    POST /{node_id}/enroll → deployment.py:806 allowlist is PENDING, ERROR, OFFLINE, DEGRADED refused: Cannot enroll node in status: online — logged twice at 13:38 today
    POST /{node_id}/reenroll → api/nodes.py:1607 requires DECOMMISSIONED unusable without decommissioning a live node first

    The node is online because its agent heartbeats. Its services have been failing for over a day. The gate reads agent reachability to answer a question about service health — the same shape as every other defect in this cluster: an instrument that cannot see the bad state reports the good one.

    I am filing that gate separately; it is a distinct defect with its own evidence and it does not block the cert fix.

    What I am fixing here

    roles/vnc includes _shared/tasks/ensure_node_tls_cert.yml instead of its own openssl block, which is what the shared task exists for. The dedicated VNC keypair below it (#13060) is untouched — that is a separate keypair for a recorded reason.

    What is still unproven

    Why the vnc role has not run on this node. Its openssl block looks correct, so a role that ran would have produced a cert. node_roles carries only slm-agent for this node while nodes.roles carries ["slm-agent", "vnc"] — two stores disagreeing about the same fact — but inventory_builder reads nodes.roles, so I cannot yet show that disagreement is the cause. Not claiming it until a provisioning run says so.

  5. mrveiss commented on Sep 8, 2026

    @mrveiss
    OwnerAuthor

    Code fix merged; the node is still down. Not closing.

    #16070 merged and verified on origin/Dev_new_gui:

    • roles/vnc/tasks/main.yml includes ensure_node_tls_cert.yml — 1 occurrence, replacing the hand-rolled openssl block
    • tls_cert_consumers_provision_it_16020_test.py carries _provisions_shared_keypair — 4 occurrences: the path-resolving classifier plus both known positives

    And the node is still failed. Live, 2026-09-08 19:28 UTC:

    nginx | failed | failed
    

    Auto-remediation has now given up seven times. That is the whole point of not closing this on the merge: the fix is correct and cannot reach the host, because nothing re-runs the vnc role on a node in this state. Closing here would tick a host-behaviour criterion with merge evidence, which is precisely the substitution this issue has already suffered once — the first fix landed in roles/nginx and roles/redis, both already correct, and was closed while the failing role went untouched.

    Filed the blocker as #16078 and recorded it as a native blocked_by edge. services/deployment.py:806 allows enrollment only from PENDING/ERROR/OFFLINE/DEGRADED, and api/nodes.py:1607 requires DECOMMISSIONED — so a node that is online because its agent heartbeats, and broken because its services are not, has no supported repair path. Two enrollment attempts today both ended Cannot enroll node in status: online.

    This issue closes when nginx on that node is running, with the host row as evidence. Not before.

  6. github-actions commented on Sep 8, 2026

    @github-actions
    Contributor

    PR #16070 (merged to Dev_new_gui) references this issue with a close keyword.

    fix(vnc): ensure the shared TLS keypair through the shared task (#16020)

    If this issue is fully resolved, close it manually. If work remains, no action is needed.

  7. added this to the Backlog milestone on Sep 12, 2026
  8. 11 remaining items

  9. mrveiss commented on Sep 14, 2026

    @mrveiss
    OwnerAuthor

    Closure audit against origin/main (vehicle merge #16702, closing PR #16522).

    AC Verdict Evidence
    Cert-ensure block exists once; no role carries a private copy met autobot-slm-backend/ansible/_shared/tasks/ensure_node_tls_cert.yml is the sole implementation; backend, frontend, nginx, redis, slm_manager, vnc all include it. git grep -rl "openssl req" .../roles/*/tasks/main.yml now returns only roles/vnc, whose block is the separate, deliberately-distinct dedicated VNC keypair (#13060), not the shared node cert
    redis provisions the cert it consumes — verified on a node whose roles are redis-only needs host evidence Code shows roles/redis/tasks/main.yml:67 includes the shared task unconditionally (not gated on other roles), which should hold on a redis-only node — but the AC's own text requires host verification, which this code-only audit does not perform
    nginx role does not start nginx before the cert exists met roles/nginx/tasks/main.yml — "Ensure the shared node TLS keypair exists (#16020)" task precedes "Ensure service is enabled and running" (state: started)
    A test asserts the consumer/provisioner sets are equal met repo_tests/tls_cert_consumers_provision_it_16020_test.py::test_every_shared_cert_consumer_ensures_the_cert_exists, ::test_nginx_ensures_the_cert_before_starting_the_service
    Failure mode is diagnosable met ensure_node_tls_cert.yml — "TLS | Confirm the shared node certificate now exists" + "TLS | Fail loudly if generation did not produce the certificate", naming the missing path and the task expected to have provided it
    Update path ensures the shared keypair on every node running a referencing service, verified by an update-all run on a slm-agent+vnc node not met autobot-slm-backend/ansible/playbooks/update-all-nodes.yml has no real reference to the vnc role (only an unrelated npm-package comment) and includes no vnc/nginx play — a node maintained by update-all whose failed services the reconciler restarts never gets the keypair. This criterion was added 2026-09-14; the fix lives on the unpushed branch issue-16020-node-tls-update-path, blocked by #16722, not yet on origin/main

    4 of 6 met; 1 needs host evidence (redis-only-node verification, not performed here); 1 not met on origin/main — its fix is on the pending branch issue-16020-node-tls-update-path, blocked by #16722. Reopening.

  10. added a commit that references this issue on Sep 19, 2026
  11. mrveiss commented on Sep 19, 2026

    @mrveiss
    OwnerAuthor

    AC verification against merged main (post #17133 vehicle merge)

    • Shared cert-ensure block, single copy — autobot-slm-backend/ansible/_shared/tasks/ensure_node_tls_cert.yml exists; confirmed 7 include sites across backend/frontend/nginx/redis/slm_manager(×2)/vnc via grep, no role carries its own openssl block.
    • redis provisions the cert, verified on a redis-only node — code fix is real: roles/redis/tasks/main.yml:65 includes the shared task with an explicit nginx cannot start on a node with no TLS cert: five roles consume the shared keypair, four provision it #16020 comment. But the AC specifically asks for verification "on a node whose roles are redis-only" — that's host behavior, not something this review can confirm from merged code. Leaving unticked.
    • nginx doesn't start before the cert exists — roles/nginx/tasks/main.yml:58 includes the shared task, :62 starts the systemd unit after.
    • Guard: the two sets are equal — repo_tests/tls_cert_consumers_provision_it_16020_test.py exists and does this sweep.
    • Diagnosable failure mode — the shared task's own ansible.builtin.fail (not a bare nginx -t exit) names the exact missing path and suggests disk space/permissions as the next check.
    • Update path ensures the cert on every relevant node, verified by an update-all run — the code-level gap is closed: repo_tests/update_all_nodes_ensures_shared_tls_cert_16020_test.py found and fixed the real structural hole (vnc/redis being skipped via tasks_from: code_only and missing _ROLE_TO_GROUPS entries, fixed via the role_vnc_active fact). But that guard's own docstring says it reads the playbook as text specifically because "parsing it as an Ansible playbook would need a live inventory + facts this repo cannot supply" — i.e., it is an explicit substitute for the literal host verification the AC asks for, not the verification itself. Leaving unticked.

    4 of 6 fully met with code evidence. The 2 remaining were already unchecked at filing and need an actual node run (redis-only, and slm-agent+vnc via update-all) to close — reopening rather than leaving this tidied by the merge.

  12. added
    blocked: needs-observationNothing but an observation on a running system is left; no session can close it
    on Oct 3, 2026
  13. mrveiss commented on Oct 3, 2026

    @mrveiss
    OwnerAuthor

    Keeping blocked: needs-observation, but recording that it is currently hiding workable code. Both halves of that sentence matter, and a sweep that acted on either one alone would get this wrong.

    A cross-issue pass flagged this issue as mislabelled on the grounds that AC6 — "The update path makes sure the shared keypair exists on every node running a service that references /etc/autobot/certs/" — is code a session can write today. The first clause is correct. The conclusion is not, because AC6 does not end there: it ends "Verified by an update-all run on a slm-agent+vnc node, whose nginx starts, or is correctly absent."

    So AC6 is not independently satisfiable — implementation and verification are welded into one criterion, and only the implementation half is reachable from a keyboard. The label governs closure, and closure still needs the run. Dropping it would move this issue into a workable queue where it would be picked up, half-delivered, and then fail to close — which is worse than the status quo, because a partial delivery closes nothing while consuming a full review.

    The real finding stands, and it is the one worth acting on: needs-observation makes an issue look like there is nothing to do, when here there is. The two remaining criteria are differently blocked and should not have been reading as one state:

    Four of six criteria are already ticked, including the guard at AC4 that stops a sixth consumer being added without the shared task. So the remaining surface is small and the implementation half of AC6 is a genuinely available piece of work for whoever takes this.

    Recommend splitting AC6 into an implementation criterion and a verification criterion, so the workable half stops being invisible behind the gated half. Not doing that unilaterally — rewording someone's acceptance criteria is a change to what "done" means here, and this issue already has a tick pattern that suggests it is being worked deliberately.

  14. mrveiss commented on Oct 3, 2026

    @mrveiss
    OwnerAuthor

    Stale premise in the 2026-10-03 09:37 comment: AC6's implementation half is already on origin/main, so there is nothing to split

    That comment says AC6's implementation is "available now" and recommends splitting it into an implementation criterion and a verification criterion. Read against origin/main, the implementation is merged:

    • playbooks/update-all-nodes.yml:1366 [PLAY 2] VNC | Ensure the shared node TLS cert, gated on role_vnc_active (not group_names, with the reason in the comment above it).
    • Also :850 backend, :1336 frontend, :2127 [PLAY 2b] Database (redis), :355 Play 1 slm_manager. Each include_tasks is _shared/tasks/ensure_node_tls_cert.yml.
    • repo_tests/update_all_nodes_ensures_shared_tls_cert_16020_test.py is the guard. It reads the playbook as text, so it does not substitute for the host run.

    So AC6 is purely observational, the same state as AC2. No criterion needs splitting.

    What discharges it. One Update-All that reaches Play 2 on a fleet containing a node with roles slm-agent + vnc. Read back on that node: /etc/autobot/certs/server-cert.pem exists, and nginx is either active or absent from the unit list.

    Dependency that is not recorded here. Every ensure task above except :355 sits in Play 2 or 2b. #12596 is open on whether Play 2 runs at all: its last comment records the detach fix as merged but no monitored Update-All showing Play 2 and the workers advancing. An Update-All that skips Play 2 would leave AC6 unmeasured, and would look like a pass on the vnc node only if the cert were already there. Check PLAY [Play 2 and PLAY RECAP in the self-update log first.

    Batches with: the same Update-All discharges #17243 (sites 3-5 are in Play 2) and the monitored-run request in #12596. AC2 (redis-only node) rides a provisioning run of a redis-only node, the topology #17331/#17242 already need.

  15. mrveiss commented on Oct 3, 2026

    @mrveiss
    OwnerAuthor

    Host observation — the consolidation has landed and holds on this node: nginx is up, nginx -t passes, and the keypair exists. The duplicated-guarantee population is gone.

    Read-only, from a live deployment at current main:

    Check Result
    Consumers of the shared cert in roles/ 6 — backend, frontend, nginx, redis, slm_manager, vnc
    Roles carrying their own openssl req block 1 (vnc, for its own separate keypair per the #13060 ruling)
    Roles including _shared/tasks/ensure_node_tls_cert.yml all 6, plus five call sites in the update-all-nodes playbook
    redis role — the role this issue found consuming without provisioning now includes the shared task before its consumer
    nginx on this node active (running); nginx -t → "configuration file test is successful"
    Shared keypair on this node both files present; openssl x509 -checkend 604800 → "Certificate will not expire"
    systemd drop-in the #1103 ExecStartPre presence tests for both cert and key are installed and in force

    What it discharges: the structural half. The invariant "any node reaching a cert consumer has a cert" is now maintained by one shared task that every consumer includes, rather than by four roles each remembering to. The specific role that broke it (redis) is fixed, and the nginx role no longer starts the daemon without first ensuring the keypair.

    What it does not discharge: the originally observed failure was a VNC node with no cert. This node has had a keypair since well before the consolidation, so the generation branch of the shared task has not been exercised here — only the "already exists, not expiring" branch. I could not determine that a node reaching a consumer with no keypair now gets one; that needs a node in that state.

    One divergence found while checking. The shared task's header states the key stays root:root 0600. On this node it is root:autobot 0640 — the file predates the consolidation, and the task gates on existence and expiry only, so it never reaches ownership on a node that already has a keypair. Not a regression introduced by the fix, but it does mean the documented ownership invariant is not enforced on any pre-existing node, only on freshly generated ones. Worth deciding whether that is intended before this issue closes.

  16. mrveiss commented on Oct 10, 2026

    @mrveiss
    OwnerAuthor

    The 3-week-old stash from issue-16020-16712-tls-restart-cap (e22f1321e0) is checked against main: every functional part is already there, in a different shape.

    • The update-all cert ensure sits in the per-role Play 1/2/2b tasks.
    • The post-generate re-check and loud failure (AC5) are present.
    • become is set in generate_self_signed_cert.yml.
    • The backend role is consolidated onto the shared task.

    One gap closed on branch issue-16020-tls-cert-stash (7ac3b7d6b3, queued): a guard ties every role that reads /etc/autobot/certs/ to an update-all task that includes the shared cert ensure, with a contrast case. A new consumer without an update-path ensure now fails.

    AC2 and AC6 are met in code and still need host evidence:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions