Repository navigation
nginx cannot start on a node with no TLS cert: five roles consume the shared keypair, four provision it #16020
Description
Activity
- added a commit that references this issue
on Sep 8, 2026 Closed on merged base, per criterion. Every quote is from
git show origin/Dev_new_gui:<path>after #16023 merged — not from the PR diff.AC1 — the cert-ensure block exists once and every consuming role includes it.
ansible/_shared/tasks/ensure_node_tls_cert.ymlis present on base. Consumers and provisioners now match exactly:roles read /etc/autobot/certs/backend,frontend,nginx,redis,slm_manager,vncensure it exists the same six Gap before:
['redis']. Gap after:[].AC2 —
redisprovisions the cert it consumes.roles/redis/tasks/main.ymlincludes the shared task, before installing Redis Stack.AC3 — nginx does not start before the certs exist.
roles/nginx/tasks/main.ymlon base: the include is at line 58,state: startedat line 65. Ordering, not merely presence — presence alone would have passed a version that ensured the cert after starting the service, which is the bug.AC4 — a test asserts the two sets are equal.
repo_tests/tls_cert_consumers_provision_it_16020_test.pycompares them as sets, with a floor on roles examined so a collapsed sweep cannot report "no gap" the way a fixed tree does. Non-vacuity is not theoretical: the guard fails on the tree that shipped.AC5 — the failure is diagnosable — NOT met, and I am not ticking it. The shared task names the missing path and the roles that provide it in its header, so a reader who reaches the file learns why. But a node that somehow still lacks the cert fails at
nginx -twith exit 1 exactly as before; nothing added a preflight that reports the missing path at the moment of failure. That criterion is unmet and stays unmet.What this does not fix
The keypair now exists on every node that reaches a consumer. A node already in the broken state is not repaired until Ansible next runs against it — this changes what provisioning produces, not what a failed node currently has. The VNC node needs a provisioning run to pick it up.
What it taught
Four roles carried hand-copied
opensslblocks and the invariant was held by each consumer separately remembering to — thevncrole's comment says so outright: "Kept here so a vnc-only node still ends up with one." A duplicated guarantee fails silently the first time it is not duplicated, on whichever node has the unlucky role set.And I nearly wrote a fifth copy.
_shared/tasks/generate_self_signed_cert.ymlalready existed from #12181, extracted fromrotate-certs.yml, with exactly one consumer. The new task delegates to it. An extraction with one consumer looks like dead weight and is actually an unadopted standard — the same shape aspre-commit-no-direct-redisusingtokenizewhile five other guards grep (#16011).Merged as #16023. Remaining live-deployment issues: #16019 (#16027 open), #16021, #16022.
Reopening the finding: the fix is correct and does not reach the failure path
Still failing on the live deployment, hours after #16023 merged and deployed. Same node, same cert, three more times today:
Sep 08 12:36:35 b9a29e04 nginx[1957982]: [emerg] cannot load certificate "/etc/autobot/certs/server-cert.pem" Sep 08 16:39:41 b9a29e04 nginx[1976462]: [emerg] cannot load certificate ... Sep 08 16:40:49 b9a29e04 nginx[1978159]: [emerg] cannot load certificate ...The deployment is current —
.deployed_commitiscbe0e646c841, identical to base — so this is not a sync lag. The fix is deployed and the failure continues.Why
ensure_node_tls_cert.ymlis included by thenginxandredisroles. The failing path runsmanage-service.yml, which has zeroroles:— it is a bare service-state play invoked byapi/orchestration.py,services/reconciler.pyandservices/playbook_executor.py.So the include never fires on the path that actually restarts nginx. I fixed provisioning; the operator's failure is in service management, and those are different playbooks.
I said at closure that "a node already in the broken state is not repaired until Ansible next runs against it" — that was right about the mechanism and wrong about the consequence. Ansible has run against it, repeatedly, via a play that cannot repair it. Waiting does not fix this node; nothing in the current design will.
Why the acceptance criteria did not catch it
Criterion 3 was "the nginx role does not start nginx before the certs its config references exist" — verified, line 58 before line 65. True, and scoped to the role. The criteria never asked "can the cert be missing on a path that starts nginx?", and
manage-servicestarts nginx while running no roles at all.That is this repository's recurring shape once more: a guard that measures the dimension the bug left. The role-level assertion holds; the defect moved to the playbook that bypasses roles.
What the fix has to be
Not a wider role include — the restart path runs no roles by construction. Either:
manage-service.ymlgains a pre-task that ensures the shared keypair when the target service is one that consumes it, or- the cert becomes a node-level invariant checked wherever a service is started, not a role-level one.
(1) is narrower and honest about what it covers; (2) is the version that stops the next consumer being missed, and matches how this issue was framed — five consumers, four provisioners.
Immediate state
This node is down and will stay down until something provisions it. The three failing restarts today were the reconciler or an operator retrying, each re-failing identically.
Reopening rather than filing a follow-up: the issue's title is exactly what is still true — nginx cannot start on a node with no TLS cert — and closing it on a role-scoped fix is what let the live failure continue unnoticed for nine hours.
Live evidence, read-only, 2026-09-08
The failure is still current and the cause is now recorded rather than inferred. From the SLM store,
services.extra_datafor nginx on the VNC node,last_checked17:47 UTC today:[emerg] cannot load certificate "/etc/autobot/certs/server-cert.pem": BIO_new_file() failed (SSL: ... No such file or directory: calling fopen(...))nginx -truns inExecStartPre, so this is not a degraded start — the unit cannot come up.Auto-remediation has given up six times in two days, each with the same terminal event:
Service nginx on VNC requires human intervention after 3 failed restart attempts 2026-09-07 20:26 · 09-08 05:44 · 09:38 · 13:47 · 16:08 · 17:45 (UTC)Fleet-wide the status picture is otherwise clean:
stopped 156 · running 77 · completed 50 · failed 4 · crash-loop 1, and zerounknown— which is separate host confirmation that the #16019/#16027 collector fix is live and working (completed 50is theactive/exitedmapping doing its job).Why the earlier close was wrong, precisely
_shared/tasks/ensure_node_tls_cert.ymlis included by exactly two roles:roles/nginx/tasks/main.yml:58 roles/redis/tasks/main.yml:67nginx on this node is not owned by
roles/nginx. It is owned byroles/vnc— androles/vnc/tasks/main.ymlstill carries its own hand-rolled openssl block at:246, one of the four duplicates the shared task's docstring was written to consolidate. It referencesserver-cert.pem(:242-254) and includes the shared task nowhere.So the fix landed in the roles that were already correct and missed the one that was failing. The consumers that reference the cert path without including the shared task are
roles/vnc,roles/backend,roles/frontend, androles/slm_manager.The second gap, which is why nobody could repair it by hand
The operator-facing repair path is
manage-service.yml, which hastasks:and noroles:— so it can restart nginx but can never establish the cert nginx needs. The rescue block added in #15879 works exactly as designed here: the[emerg]line above reached the controller log instead of staying on the node. It made the failure legible; it cannot make it repairable.And the two paths that could re-run provisioning both refuse this node:
path gate result POST /{node_id}/enroll→deployment.py:806allowlist is PENDING, ERROR, OFFLINE, DEGRADEDrefused: Cannot enroll node in status: online— logged twice at 13:38 todayPOST /{node_id}/reenroll→api/nodes.py:1607requires DECOMMISSIONEDunusable without decommissioning a live node first The node is
onlinebecause its agent heartbeats. Its services have been failing for over a day. The gate reads agent reachability to answer a question about service health — the same shape as every other defect in this cluster: an instrument that cannot see the bad state reports the good one.I am filing that gate separately; it is a distinct defect with its own evidence and it does not block the cert fix.
What I am fixing here
roles/vncincludes_shared/tasks/ensure_node_tls_cert.ymlinstead of its own openssl block, which is what the shared task exists for. The dedicated VNC keypair below it (#13060) is untouched — that is a separate keypair for a recorded reason.What is still unproven
Why the vnc role has not run on this node. Its openssl block looks correct, so a role that ran would have produced a cert.
node_rolescarries onlyslm-agentfor this node whilenodes.rolescarries["slm-agent", "vnc"]— two stores disagreeing about the same fact — butinventory_builderreadsnodes.roles, so I cannot yet show that disagreement is the cause. Not claiming it until a provisioning run says so.- added a commit that references this issue
on Sep 8, 2026 Code fix merged; the node is still down. Not closing.
#16070 merged and verified on
origin/Dev_new_gui:roles/vnc/tasks/main.ymlincludesensure_node_tls_cert.yml— 1 occurrence, replacing the hand-rolled openssl blocktls_cert_consumers_provision_it_16020_test.pycarries_provisions_shared_keypair— 4 occurrences: the path-resolving classifier plus both known positives
And the node is still failed. Live, 2026-09-08 19:28 UTC:
nginx | failed | failedAuto-remediation has now given up seven times. That is the whole point of not closing this on the merge: the fix is correct and cannot reach the host, because nothing re-runs the
vncrole on a node in this state. Closing here would tick a host-behaviour criterion with merge evidence, which is precisely the substitution this issue has already suffered once — the first fix landed inroles/nginxandroles/redis, both already correct, and was closed while the failing role went untouched.Filed the blocker as #16078 and recorded it as a native
blocked_byedge.services/deployment.py:806allows enrollment only fromPENDING/ERROR/OFFLINE/DEGRADED, andapi/nodes.py:1607requiresDECOMMISSIONED— so a node that isonlinebecause its agent heartbeats, and broken because its services are not, has no supported repair path. Two enrollment attempts today both endedCannot enroll node in status: online.This issue closes when nginx on that node is running, with the host row as evidence. Not before.
11 remaining items
Closure audit against
origin/main(vehicle merge #16702, closing PR #16522).AC Verdict Evidence Cert-ensure block exists once; no role carries a private copy met autobot-slm-backend/ansible/_shared/tasks/ensure_node_tls_cert.ymlis the sole implementation;backend,frontend,nginx,redis,slm_manager,vncall include it.git grep -rl "openssl req" .../roles/*/tasks/main.ymlnow returns onlyroles/vnc, whose block is the separate, deliberately-distinct dedicated VNC keypair (#13060), not the shared node certredisprovisions the cert it consumes — verified on a node whose roles are redis-onlyneeds host evidence Code shows roles/redis/tasks/main.yml:67includes the shared task unconditionally (not gated on other roles), which should hold on a redis-only node — but the AC's own text requires host verification, which this code-only audit does not performnginxrole does not start nginx before the cert existsmet roles/nginx/tasks/main.yml— "Ensure the shared node TLS keypair exists (#16020)" task precedes "Ensure service is enabled and running" (state: started)A test asserts the consumer/provisioner sets are equal met repo_tests/tls_cert_consumers_provision_it_16020_test.py::test_every_shared_cert_consumer_ensures_the_cert_exists,::test_nginx_ensures_the_cert_before_starting_the_serviceFailure mode is diagnosable met ensure_node_tls_cert.yml— "TLS | Confirm the shared node certificate now exists" + "TLS | Fail loudly if generation did not produce the certificate", naming the missing path and the task expected to have provided itUpdate path ensures the shared keypair on every node running a referencing service, verified by an update-all run on a slm-agent+vncnodenot met autobot-slm-backend/ansible/playbooks/update-all-nodes.ymlhas no real reference to thevncrole (only an unrelated npm-package comment) and includes novnc/nginxplay — a node maintained by update-all whose failed services the reconciler restarts never gets the keypair. This criterion was added 2026-09-14; the fix lives on the unpushed branchissue-16020-node-tls-update-path, blocked by #16722, not yet onorigin/main4 of 6 met; 1 needs host evidence (redis-only-node verification, not performed here); 1 not met on
origin/main— its fix is on the pending branchissue-16020-node-tls-update-path, blocked by #16722. Reopening.- added a commit that references this issue
on Sep 19, 2026 AC verification against merged main (post #17133 vehicle merge)
- Shared cert-ensure block, single copy —
autobot-slm-backend/ansible/_shared/tasks/ensure_node_tls_cert.ymlexists; confirmed 7 include sites across backend/frontend/nginx/redis/slm_manager(×2)/vnc via grep, no role carries its own openssl block. - redis provisions the cert, verified on a redis-only node — code fix is real:
roles/redis/tasks/main.yml:65includes the shared task with an explicit nginx cannot start on a node with no TLS cert: five roles consume the shared keypair, four provision it #16020 comment. But the AC specifically asks for verification "on a node whose roles are redis-only" — that's host behavior, not something this review can confirm from merged code. Leaving unticked. - nginx doesn't start before the cert exists —
roles/nginx/tasks/main.yml:58includes the shared task,:62starts the systemd unit after. - Guard: the two sets are equal —
repo_tests/tls_cert_consumers_provision_it_16020_test.pyexists and does this sweep. - Diagnosable failure mode — the shared task's own
ansible.builtin.fail(not a barenginx -texit) names the exact missing path and suggests disk space/permissions as the next check. - Update path ensures the cert on every relevant node, verified by an update-all run — the code-level gap is closed:
repo_tests/update_all_nodes_ensures_shared_tls_cert_16020_test.pyfound and fixed the real structural hole (vnc/redis being skipped viatasks_from: code_onlyand missing_ROLE_TO_GROUPSentries, fixed via therole_vnc_activefact). But that guard's own docstring says it reads the playbook as text specifically because "parsing it as an Ansible playbook would need a live inventory + facts this repo cannot supply" — i.e., it is an explicit substitute for the literal host verification the AC asks for, not the verification itself. Leaving unticked.
4 of 6 fully met with code evidence. The 2 remaining were already unchecked at filing and need an actual node run (redis-only, and slm-agent+vnc via update-all) to close — reopening rather than leaving this tidied by the merge.
- Shared cert-ensure block, single copy —
- addedblocked: needs-observationNothing but an observation on a running system is left; no session can close itNothing but an observation on a running system is left; no session can close it
on Oct 3, 2026 Keeping
blocked: needs-observation, but recording that it is currently hiding workable code. Both halves of that sentence matter, and a sweep that acted on either one alone would get this wrong.A cross-issue pass flagged this issue as mislabelled on the grounds that AC6 — "The update path makes sure the shared keypair exists on every node running a service that references
/etc/autobot/certs/" — is code a session can write today. The first clause is correct. The conclusion is not, because AC6 does not end there: it ends "Verified by an update-all run on aslm-agent+vncnode, whose nginx starts, or is correctly absent."So AC6 is not independently satisfiable — implementation and verification are welded into one criterion, and only the implementation half is reachable from a keyboard. The label governs closure, and closure still needs the run. Dropping it would move this issue into a workable queue where it would be picked up, half-delivered, and then fail to close — which is worse than the status quo, because a partial delivery closes nothing while consuming a full review.
The real finding stands, and it is the one worth acting on:
needs-observationmakes an issue look like there is nothing to do, when here there is. The two remaining criteria are differently blocked and should not have been reading as one state:- AC2 — "
redisprovisions the cert it consumes — verified on a node whose roles are redis-only" — is purely observational. Nothing can be written toward it; it needs a role-separated topology to exist, which is what the multi-host deploy (CRITICAL(deploy): the filtered backend manifest bakes a controller-only constraint path, so provisioning still aborts on every non-manager node #17331, deploy: five more tasks read code_source on hosts that do not have it — three are in the builtin updater (#17242 class) #17243, CRITICAL(deploy): backend provisioning aborts on every non-manager host — a task executes a script out of code_source, which exists only on the controller (rc=127) #17242) produces. - AC6 — implementation available now, verification gated behind an update-all run on an
slm-agent+vncnode. Note this is a different observation from AC2's, and the cheaper one.
Four of six criteria are already ticked, including the guard at AC4 that stops a sixth consumer being added without the shared task. So the remaining surface is small and the implementation half of AC6 is a genuinely available piece of work for whoever takes this.
Recommend splitting AC6 into an implementation criterion and a verification criterion, so the workable half stops being invisible behind the gated half. Not doing that unilaterally — rewording someone's acceptance criteria is a change to what "done" means here, and this issue already has a tick pattern that suggests it is being worked deliberately.
- AC2 — "
Stale premise in the 2026-10-03 09:37 comment: AC6's implementation half is already on
origin/main, so there is nothing to splitThat comment says AC6's implementation is "available now" and recommends splitting it into an implementation criterion and a verification criterion. Read against
origin/main, the implementation is merged:playbooks/update-all-nodes.yml:1366[PLAY 2] VNC | Ensure the shared node TLS cert, gated onrole_vnc_active(notgroup_names, with the reason in the comment above it).- Also
:850backend,:1336frontend,:2127[PLAY 2b] Database(redis),:355Play 1 slm_manager. Eachinclude_tasksis_shared/tasks/ensure_node_tls_cert.yml. repo_tests/update_all_nodes_ensures_shared_tls_cert_16020_test.pyis the guard. It reads the playbook as text, so it does not substitute for the host run.
So AC6 is purely observational, the same state as AC2. No criterion needs splitting.
What discharges it. One Update-All that reaches Play 2 on a fleet containing a node with roles
slm-agent+vnc. Read back on that node:/etc/autobot/certs/server-cert.pemexists, and nginx is either active or absent from the unit list.Dependency that is not recorded here. Every ensure task above except
:355sits in Play 2 or 2b. #12596 is open on whether Play 2 runs at all: its last comment records the detach fix as merged but no monitored Update-All showing Play 2 and the workers advancing. An Update-All that skips Play 2 would leave AC6 unmeasured, and would look like a pass on the vnc node only if the cert were already there. CheckPLAY [Play 2andPLAY RECAPin the self-update log first.Batches with: the same Update-All discharges #17243 (sites 3-5 are in Play 2) and the monitored-run request in #12596. AC2 (redis-only node) rides a provisioning run of a redis-only node, the topology #17331/#17242 already need.
Host observation — the consolidation has landed and holds on this node: nginx is up,
nginx -tpasses, and the keypair exists. The duplicated-guarantee population is gone.Read-only, from a live deployment at current
main:Check Result Consumers of the shared cert in roles/6 — backend,frontend,nginx,redis,slm_manager,vncRoles carrying their own openssl reqblock1 ( vnc, for its own separate keypair per the #13060 ruling)Roles including _shared/tasks/ensure_node_tls_cert.ymlall 6, plus five call sites in the update-all-nodes playbook redisrole — the role this issue found consuming without provisioningnow includes the shared task before its consumer nginx on this node active (running);nginx -t→ "configuration file test is successful"Shared keypair on this node both files present; openssl x509 -checkend 604800→ "Certificate will not expire"systemd drop-in the #1103 ExecStartPrepresence tests for both cert and key are installed and in forceWhat it discharges: the structural half. The invariant "any node reaching a cert consumer has a cert" is now maintained by one shared task that every consumer includes, rather than by four roles each remembering to. The specific role that broke it (
redis) is fixed, and thenginxrole no longer starts the daemon without first ensuring the keypair.What it does not discharge: the originally observed failure was a VNC node with no cert. This node has had a keypair since well before the consolidation, so the generation branch of the shared task has not been exercised here — only the "already exists, not expiring" branch. I could not determine that a node reaching a consumer with no keypair now gets one; that needs a node in that state.
One divergence found while checking. The shared task's header states the key stays
root:root 0600. On this node it isroot:autobot 0640— the file predates the consolidation, and the task gates on existence and expiry only, so it never reaches ownership on a node that already has a keypair. Not a regression introduced by the fix, but it does mean the documented ownership invariant is not enforced on any pre-existing node, only on freshly generated ones. Worth deciding whether that is intended before this issue closes.The 3-week-old stash from
issue-16020-16712-tls-restart-cap(e22f1321e0) is checked againstmain: every functional part is already there, in a different shape.- The update-all cert ensure sits in the per-role Play 1/2/2b tasks.
- The post-generate re-check and loud failure (AC5) are present.
becomeis set ingenerate_self_signed_cert.yml.- The backend role is consolidated onto the shared task.
One gap closed on branch
issue-16020-tls-cert-stash(7ac3b7d6b3, queued): a guard ties every role that reads/etc/autobot/certs/to an update-all task that includes the shared cert ensure, with a contrast case. A new consumer without an update-path ensure now fails.AC2 and AC6 are met in code and still need host evidence:
- AC2: a redis-only node's
/etc/autobot/certs/server-cert.pem. - AC6: one Update-All reaching Play 2 on an
slm-agent+vncnode. The run's log must showPLAY [Play 2andPLAY RECAP, given deploy: Update-All STILL skips Play 2 — #12567 detach activates but the transient scope is killed by the Play-1 SLM restart (incomplete #12425) #12596.
Symptom
nginx is failing to start on the VNC node and has been since at least 2026-09-07 23:18 — six consecutive failed restarts, each blocking the
manage-serviceplaybook.The nginx config references a certificate the node does not have.
ExecStartPre=nginx -tfails, so the unit can never come up, and every restart attempt re-fails identically.The defect is a population, not an instance
Counted across
autobot-slm-backend/ansible/roles/:/etc/autobot/certs/server-cert.pembackend,frontend,redis,slm_manager,vncopenssl req)backend,frontend,slm_manager,vncredisconsumes the shared keypair and never provisions it. It is the same latent failure as the nginx one, waiting for a node whose role set does not happen to include a provisioner.And the
nginxrole itself does exactly three things — refresh apt, install nginx, start the service. It starts a daemon whose config may reference a cert it never checks for.Why it survived
Four roles each carry their own copy of the same
openssl reqblock. Thevncrole's own comment records the pattern and its reason:So the invariant "any node reaching a cert consumer has a cert" is maintained by each consumer separately remembering to, with nothing enforcing it. It holds for as long as nobody adds a consumer without a copy. Someone added
redis.A duplicated guarantee is one that fails silently the first time it is not duplicated, and it fails on whichever node has the unlucky role set — which is why this shows up on the VNC node and not on the control node.
Fix direction
One shared task file —
roles/_shared/tasks/ensure_node_tls_cert.yml, alongside the six that already live there — included by every consumer, replacing four hand-copied blocks with one.nginxincludes it before starting the service;redisgets it for the first time.Acceptance criteria
redisprovisions the cert it consumes — verified on a node whose roles are redis-onlynginxrole does not start nginx before the certs its config references exist/etc/autobot/certs/includes the shared task. A guard over the roles tree, so a sixth consumer cannot be added without onenginx -texit 1/etc/autobot/certs/. Today no role that update-all runs on this node does so. Verified by an update-all run on aslm-agent+vncnode, whose nginx starts, or is correctly absent.The fourth criterion is what stops this recurring. The current invariant lives in a comment; an invariant a sweep can read is a control, and prose is not.
Blocks the operator's VNC node. Related: #16022 (setup cannot get past provision) may be this same failure surfacing in the provisioning path — filed separately because the log lines proving the nginx failure all come from the service-management playbook, not a provisioning one. Also #16019, #16021.