Repository navigation
security(chromadb): the vector store binds every interface in front of four unpatchable advisories, one a pre-auth RCE #15317
Description
Activity
- addedneeds-decisionBlocked on an owner decision; options and a recommendation are on the issueBlocked on an owner decision; options and a recommendation are on the issue
on Aug 30, 2026 Owner decision, 2026-09-18: as proposed. The owner plans to cluster ChromaDB later, which decides the design.
The deciding fact is that one of the four advisories is pre-authentication. A pre-auth flaw is reached before any credential is checked, so application-level authentication does not mitigate it. Only a control that gates the connection before it reaches ChromaDB's code does. That rules the options in and out:
- Loopback-only is a dead end — clustering needs node-to-node reachability by definition.
0.0.0.0is never right, and two templates hardcode it against the no-hardcode rule.- The private cluster interface is the right default: it scales to N nodes and removes public exposure.
- Transport-level mutual TLS is the only complete fix for the pre-auth advisory.
So, in two parts:
- This issue — now. Split
chromadb_hostinto a bind variable and a dial variable. That split is the root cause named above: one variable doing two incompatible jobs is why the bind was left wide, since narrowing it broke the clients. Default the bind to the private cluster interface. Remove the two hardcoded--host 0.0.0.0templates. This is the same bind/dial conflation test(ci): 28 tests under autobot-infrastructure/shared/tests and libs/ run in no workflow, and ci.yml's integration step points at a path that does not exist #15051 already fixed for the backend, so there is a pattern to follow. PR security(infra): stop binding ChromaDB to every interface by default (#15317) #16882 is the in-flight work here. - mTLS — with the clustering work, filed separately and linked below. Application-level auth is still worth adding as defence in depth for the other three advisories, but it does not address this one.
One thing that could not be determined, and the fix should state rather than assume: whether each deployment actually has a genuinely private network segment. "Private interface" is only as safe as the network it sits on.
- removedneeds-decisionBlocked on an owner decision; options and a recommendation are on the issueBlocked on an owner decision; options and a recommendation are on the issue
on Sep 18, 2026 The mTLS half is filed as #16945, scheduled with the clustering work.
- added 7 commits that reference this issue
on Sep 18, 2026 Per-AC closure pass on
origin/mainaf1c060 (read-only; no verdict on closing)Result:
- Met: AC1, AC4, AC5.
- Partly met: AC2.
- Needs an owner call: AC3. The shipped default is loopback, which is Option A. The owner ruling of 2026-09-18 (comment 5725701035) chose the private cluster interface and called loopback "a dead end".
- Not measurable from the tree: AC6.
Commits naming #15317, oldest first (
git log --reverse --grep='#15317' origin/main):12d810f6c2026-09-17 chore: claim worktree for security(chromadb): the vector store binds every interface in front of four unpatchable advisories, one a pre-auth RCE #153171c2a9778c2026-09-17 security(infra): stop binding ChromaDB to every interface by default5170cca712026-09-18 fix(security): close the chromadb bind guard's two review gaps4e7ddf8ed2026-09-18 fix(security): close 3 nits from the chromadb bind guard's re-review9410b8c362026-09-18 fix(security): restore empty-bind alternation and catch chromadb host networking7d5b53e412026-09-19 fix(repo_tests): record security(chromadb): the vector store binds every interface in front of four unpatchable advisories, one a pre-auth RCE #15317's .service/.service.j2/docker/*.yml globs (kb_read_visibility_guard_test red on main: claude_memory_importer.import_memory_file has no ownership filter #17124)1b3ffd58e2026-09-22 chore(vehicle): land 4 small branches (… chore(vehicle): land 4 small branches (#16249, #16394, CodeQL fixes, dependabot uv) #17116). Mentions this issue only.
Note the order: the default was chosen in
1c2a9778con 09-17, before the owner ruling (09-18 05:40). The three 09-18 follow-ups changed the guard, not the default.# AC Status Evidence 1 Bind and dial are separate variables; no template hardcodes either Met (one observation) Quotes 1a–1c below. 2 All three templates take the bind from configuration; no --host 0.0.0.0literal anywherePartly met Quotes 2a–2c below. 3 Chosen default applied consistently across all three, decision recorded inline Consistent and recorded, but not the default the owner chose Quotes 3a–3c below. 4 A guard stops any ChromaDB template carrying a hardcoded bind literal Met Quotes 4a–4c below. 5 Multi-node docs state the bind variable and a subnet-scoped firewall rule, never allow-from-anywhere Met Quotes 5a–5b below. 6 Verified on a deployed host: listens only on the intended interface, consumers still reach it Not measurable from the tree Needs a host observation, e.g. ss -ltnpon the ChromaDB node plus one client query from each consumer role. None was made here.AC1 quotes:
- 1a. The bind variable is
chromadb_bind_host: "127.0.0.1"(ansible/inventory/group_vars/all.yml:441). The dial variable ischromadb_host(roles/ai-stack/defaults/main.yml:30), consumed asCHROMADB_HOST={{ chromadb_host }}(ai-stack.env.j2:22). Neither service template hardcodes either value. - 1b. The dial default is still
"0.0.0.0". The inline note atdefaults/main.yml:26-29reads: "DIAL address only now … Left at its pre-existing value rather than changed here: this variable's actual consumers were not fully traced." So that value is not a bind violation. - 1c. The issue body itself says
0.0.0.0"is not a meaningful destination" for a dial address. That half of the conflation is recorded but left untraced. The backend dials through its ownbackend_chromadb_host(roles/backend/tasks/main.yml:76,defaults/main.yml:149), notchromadb_host.
AC2 quotes:
- 2a.
roles/ai-stack/templates/autobot-chromadb.service.j2:54:--host {{ chromadb_bind_host }}. - 2b.
roles/redis/templates/autobot-chromadb.service.j2:52:--host {{ chromadb_bind_host }}. - 2c.
autobot-infrastructure/autobot-database/templates/autobot-chromadb.service:39is still a literal:--host 127.0.0.1. It is a static non-Jinja file, and:35-36calls it a "static reference copy … kept in step". The0.0.0.0half of the criterion is met. The "from configuration" half is not met for this file. - Sweep:
git grep -E -- '--host[ =]+0\.0\.0\.0'on main hits only the changelog and the guard's own synthetic fixtures.
AC3 quotes:
- 3a. All three files use loopback.
- 3b. The decision is recorded inline at
group_vars/all.yml:433-440: "Loopback by default (Option A): a multi-node install must set this explicitly and add a subnet-scoped firewall rule". It is echoed atredis/…service.j2:44-48. - 3c. The owner ruling reads: "Loopback-only is a dead end — clustering needs node-to-node reachability … The private cluster interface is the right default … Default the bind to the private cluster interface." The tree does not implement that default. Whether loopback-plus-explicit-override is acceptable for now is the owner's call.
AC4 quotes:
- 4a. The guard is
repo_tests/chromadb_bind_not_hardcoded_15317_test.py, with_LAUNCH_SITE_PATTERNS = ("*.service", "*.service.j2", "*.sh")(:87). - 4b. It has a coverage floor (
test_the_sweep_reaches_the_known_chroma_launch_sites,:326) and the assertiontest_no_chroma_launch_site_hardcodes_a_wide_bind(:332). - 4c. It also has override, compose-port and host-networking checks (
:374,:417,:430), each with a synthetic negative control (:442,:468,:477,:506,:523). CI runsrepo_testsin.github/workflows/ci.yml:344("Run unit tests — backend, shared, tts-worker, repo_tests"). One limit: the guard targets wildcard binds. A literal non-wildcard like 2c's127.0.0.1passes it by design.
AC5 quotes:
- 5a.
docs/architecture/NETWORK_TOPOLOGY.md:198-213("ChromaDB Bind Address (security(chromadb): the vector store binds every interface in front of four unpatchable advisories, one a pre-auth RCE #15317)") nameschromadb_bind_host, separates it from the dialchromadb_host, and says a multi-node install must set it explicitly "and add a UFW rule scoped to the calling node's IP and port 8100". - 5b. It warns that leaving
0.0.0.0reopens the exposure. It contains no allow-from-anywhere rule.
Read vs inferred:
- Read: the issue body and both comments; the lines quoted above from all three unit files, both defaults files,
ai-stack.env.j2, the guard's test list and the docs section; and therepo_testsstep inci.yml. - Not run: the guard. I also did not read its full body; its coverage is taken from test names and the lines cited.
- Inferred: that nothing else consumes
chromadb_hostas a bind. The grep was overautobot-slm-backend/ansibleforchromadb_host:only.
Undetermined:
- Whether each deployment has a genuinely private segment. The owner ruling asked the fix to state this. I found no such statement in the tree, but searched only the files above.
- What actually consumes the
0.0.0.0dial default.
Every criterion was reached; none was skipped.
Generated by Claude Code
Verdict on the closure pass above: stays open on 2c.
autobot-infrastructure/autobot-database/templates/autobot-chromadb.service:39still carries a literal--host 127.0.0.1in a static non-Jinja reference copy; the0.0.0.0half of the criterion is met, thefrom configurationhalf is not for that file. Either the reference copy is templated or the criterion states that a static reference copy is exempt — one line of work either way. Not closed.
What
ChromaDB binds every interface in front of four advisories that have no upstream fix, one of which is a pre-authentication code injection with a published proof-of-concept. The host firewall is currently the only control standing in front of it.
Found while working #15052, which asked for a version bump. There is none to make:
chromadb 1.5.9is the latest published release, and OSV reports no fixed version for any of the four (details and the advisory table are on #15052). Since the vulnerabilities cannot be patched, exposure is the only remaining control — which makes the bind address the thing that matters.The bind
Three service templates start the server, and they disagree:
autobot-slm-backend/ansible/roles/ai-stack/templates/autobot-chromadb.service.j2:54{{ chromadb_host }}— parameterised, but see the default belowautobot-slm-backend/ansible/roles/redis/templates/autobot-chromadb.service.j2:47--host 0.0.0.0autobot-infrastructure/autobot-database/templates/autobot-chromadb.service:37--host 0.0.0.0Two hardcode the value, which is also a straight violation of the "never hardcode" rule. And the one that is parameterised does not help, because
autobot-slm-backend/ansible/roles/ai-stack/defaults/main.yml:26sets:So all three bind all interfaces.
Root cause: one variable is doing two incompatible jobs
chromadb_hostis used as the bind address in the unit'sExecStart, and as the dial address handed to clients (autobot-slm-backend/ansible/roles/ai-stack/templates/ai-stack.env.j2:22→CHROMADB_HOST). Those are different values with different safe defaults:0.0.0.0is not a meaningful destinationBecause they share one variable, narrowing the bind would break the clients, which is very likely why it was left wide. This is the same conflation #15051 just fixed on the backend, where
get_backend_config()exposed onlyserver_host/server_port(the bind) whileservice_discovery.pywas readinghost/port(the dial) and logging an error on every process start. Same defect, different service.Current exposure
On the live host, checked directly:
So there is no live incident here. The problem is that the protection is one deep: token auth does not apply to a pre-authentication vulnerability, so the firewall is the entire defence. A rule added for debugging, a host brought up under a different profile, or a container network that bypasses the host firewall turns an unpatchable critical RCE into a directly reachable one, with nothing behind it.
The decision this needs
Splitting bind from dial is not controversial and should happen regardless. What needs a decision is the default bind, because AutoBot scales from Docker to a single VM to any number of nodes (#15194), and the right default differs:
I recommend A. It is the only one that makes the safe case the default, and the breakage it can cause is loud and immediate (a multi-node install cannot reach its vector store) rather than silent — which is the right failure direction for a security default. The migration is one variable in the inventory.
Acceptance criteria
--host 0.0.0.0literal anywhereRefs