Skip to content

security(chromadb): the vector store binds every interface in front of four unpatchable advisories, one a pre-auth RCE #15317

Description

@mrveiss

What

ChromaDB binds every interface in front of four advisories that have no upstream fix, one of which is a pre-authentication code injection with a published proof-of-concept. The host firewall is currently the only control standing in front of it.

Found while working #15052, which asked for a version bump. There is none to make: chromadb 1.5.9 is the latest published release, and OSV reports no fixed version for any of the four (details and the advisory table are on #15052). Since the vulnerabilities cannot be patched, exposure is the only remaining control — which makes the bind address the thing that matters.

The bind

Three service templates start the server, and they disagree:

Template Bind
autobot-slm-backend/ansible/roles/ai-stack/templates/autobot-chromadb.service.j2:54 {{ chromadb_host }} — parameterised, but see the default below
autobot-slm-backend/ansible/roles/redis/templates/autobot-chromadb.service.j2:47 hardcoded --host 0.0.0.0
autobot-infrastructure/autobot-database/templates/autobot-chromadb.service:37 hardcoded --host 0.0.0.0

Two hardcode the value, which is also a straight violation of the "never hardcode" rule. And the one that is parameterised does not help, because autobot-slm-backend/ansible/roles/ai-stack/defaults/main.yml:26 sets:

chromadb_host: "0.0.0.0"

So all three bind all interfaces.

Root cause: one variable is doing two incompatible jobs

chromadb_host is used as the bind address in the unit's ExecStart, and as the dial address handed to clients (autobot-slm-backend/ansible/roles/ai-stack/templates/ai-stack.env.j2:22 → CHROMADB_HOST). Those are different values with different safe defaults:

  • a bind address should be the narrowest interface that still serves its consumers — loopback when they are co-located
  • a dial address must be one a client can actually reach, and 0.0.0.0 is not a meaningful destination

Because they share one variable, narrowing the bind would break the clients, which is very likely why it was left wide. This is the same conflation #15051 just fixed on the backend, where get_backend_config() exposed only server_host/server_port (the bind) while service_discovery.py was reading host/port (the dial) and logging an error on every process start. Same defect, different service.

Current exposure

On the live host, checked directly:

So there is no live incident here. The problem is that the protection is one deep: token auth does not apply to a pre-authentication vulnerability, so the firewall is the entire defence. A rule added for debugging, a host brought up under a different profile, or a container network that bypasses the host firewall turns an unpatchable critical RCE into a directly reachable one, with nothing behind it.

The decision this needs

Splitting bind from dial is not controversial and should happen regardless. What needs a decision is the default bind, because AutoBot scales from Docker to a single VM to any number of nodes (#15194), and the right default differs:

Option Consequence
A (recommended) Default bind to loopback; multi-node deployments set the bind explicitly and add a subnet-scoped firewall rule Safe by default. Existing multi-node installs that relied on the wide bind need the explicit setting at next update, or ChromaDB becomes unreachable for them.
B Keep the wide bind as the default; add the split and document the risk Nothing breaks. The exposure remains exactly as it is, and the documentation is the only control added — which is not a control.
C Bind to the internal-interface address by default, loopback when the consumer is co-located No breakage and a real narrowing, but the "is the consumer co-located" test has to be derived during deployment, which is more machinery and more ways to be wrong.

I recommend A. It is the only one that makes the safe case the default, and the breakage it can cause is loud and immediate (a multi-node install cannot reach its vector store) rather than silent — which is the right failure direction for a security default. The migration is one variable in the inventory.

Acceptance criteria

  • Bind address and client dial address are separate variables; no template hardcodes either
  • All three templates take the bind from configuration — no --host 0.0.0.0 literal anywhere
  • The chosen default is applied consistently across all three, with the decision recorded inline
  • A guard asserts no ChromaDB service template contains a hardcoded bind literal, so this cannot regress
  • Multi-node deployment documentation states the bind variable and the subnet-scoped firewall rule it requires — never an "allow from anywhere" rule
  • Verified on a deployed host: the service listens only on the intended interface, and its consumers still reach it

Refs

Activity

  1. added this to the v0.9.0 milestone on Sep 12, 2026
  2. mrveiss commented on Sep 18, 2026

    @mrveiss
    OwnerAuthor

    Owner decision, 2026-09-18: as proposed. The owner plans to cluster ChromaDB later, which decides the design.

    The deciding fact is that one of the four advisories is pre-authentication. A pre-auth flaw is reached before any credential is checked, so application-level authentication does not mitigate it. Only a control that gates the connection before it reaches ChromaDB's code does. That rules the options in and out:

    • Loopback-only is a dead end — clustering needs node-to-node reachability by definition.
    • 0.0.0.0 is never right, and two templates hardcode it against the no-hardcode rule.
    • The private cluster interface is the right default: it scales to N nodes and removes public exposure.
    • Transport-level mutual TLS is the only complete fix for the pre-auth advisory.

    So, in two parts:

    1. This issue — now. Split chromadb_host into a bind variable and a dial variable. That split is the root cause named above: one variable doing two incompatible jobs is why the bind was left wide, since narrowing it broke the clients. Default the bind to the private cluster interface. Remove the two hardcoded --host 0.0.0.0 templates. This is the same bind/dial conflation test(ci): 28 tests under autobot-infrastructure/shared/tests and libs/ run in no workflow, and ci.yml's integration step points at a path that does not exist #15051 already fixed for the backend, so there is a pattern to follow. PR security(infra): stop binding ChromaDB to every interface by default (#15317) #16882 is the in-flight work here.
    2. mTLS — with the clustering work, filed separately and linked below. Application-level auth is still worth adding as defence in depth for the other three advisories, but it does not address this one.

    One thing that could not be determined, and the fix should state rather than assume: whether each deployment actually has a genuinely private network segment. "Private interface" is only as safe as the network it sits on.

  3. removed
    needs-decisionBlocked on an owner decision; options and a recommendation are on the issue
    on Sep 18, 2026
  4. mrveiss commented on Sep 18, 2026

    @mrveiss
    OwnerAuthor

    The mTLS half is filed as #16945, scheduled with the clustering work.

  5. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    Per-AC closure pass on origin/main af1c060 (read-only; no verdict on closing)

    Result:

    • Met: AC1, AC4, AC5.
    • Partly met: AC2.
    • Needs an owner call: AC3. The shipped default is loopback, which is Option A. The owner ruling of 2026-09-18 (comment 5725701035) chose the private cluster interface and called loopback "a dead end".
    • Not measurable from the tree: AC6.

    Commits naming #15317, oldest first (git log --reverse --grep='#15317' origin/main):

    1. 12d810f6c 2026-09-17 chore: claim worktree for security(chromadb): the vector store binds every interface in front of four unpatchable advisories, one a pre-auth RCE #15317
    2. 1c2a9778c 2026-09-17 security(infra): stop binding ChromaDB to every interface by default
    3. 5170cca71 2026-09-18 fix(security): close the chromadb bind guard's two review gaps
    4. 4e7ddf8ed 2026-09-18 fix(security): close 3 nits from the chromadb bind guard's re-review
    5. 9410b8c36 2026-09-18 fix(security): restore empty-bind alternation and catch chromadb host networking
    6. 7d5b53e41 2026-09-19 fix(repo_tests): record security(chromadb): the vector store binds every interface in front of four unpatchable advisories, one a pre-auth RCE #15317's .service/.service.j2/docker/*.yml globs (kb_read_visibility_guard_test red on main: claude_memory_importer.import_memory_file has no ownership filter #17124)
    7. 1b3ffd58e 2026-09-22 chore(vehicle): land 4 small branches (… chore(vehicle): land 4 small branches (#16249, #16394, CodeQL fixes, dependabot uv) #17116). Mentions this issue only.

    Note the order: the default was chosen in 1c2a9778c on 09-17, before the owner ruling (09-18 05:40). The three 09-18 follow-ups changed the guard, not the default.

    # AC Status Evidence
    1 Bind and dial are separate variables; no template hardcodes either Met (one observation) Quotes 1a–1c below.
    2 All three templates take the bind from configuration; no --host 0.0.0.0 literal anywhere Partly met Quotes 2a–2c below.
    3 Chosen default applied consistently across all three, decision recorded inline Consistent and recorded, but not the default the owner chose Quotes 3a–3c below.
    4 A guard stops any ChromaDB template carrying a hardcoded bind literal Met Quotes 4a–4c below.
    5 Multi-node docs state the bind variable and a subnet-scoped firewall rule, never allow-from-anywhere Met Quotes 5a–5b below.
    6 Verified on a deployed host: listens only on the intended interface, consumers still reach it Not measurable from the tree Needs a host observation, e.g. ss -ltnp on the ChromaDB node plus one client query from each consumer role. None was made here.

    AC1 quotes:

    • 1a. The bind variable is chromadb_bind_host: "127.0.0.1" (ansible/inventory/group_vars/all.yml:441). The dial variable is chromadb_host (roles/ai-stack/defaults/main.yml:30), consumed as CHROMADB_HOST={{ chromadb_host }} (ai-stack.env.j2:22). Neither service template hardcodes either value.
    • 1b. The dial default is still "0.0.0.0". The inline note at defaults/main.yml:26-29 reads: "DIAL address only now … Left at its pre-existing value rather than changed here: this variable's actual consumers were not fully traced." So that value is not a bind violation.
    • 1c. The issue body itself says 0.0.0.0 "is not a meaningful destination" for a dial address. That half of the conflation is recorded but left untraced. The backend dials through its own backend_chromadb_host (roles/backend/tasks/main.yml:76, defaults/main.yml:149), not chromadb_host.

    AC2 quotes:

    • 2a. roles/ai-stack/templates/autobot-chromadb.service.j2:54: --host {{ chromadb_bind_host }}.
    • 2b. roles/redis/templates/autobot-chromadb.service.j2:52: --host {{ chromadb_bind_host }}.
    • 2c. autobot-infrastructure/autobot-database/templates/autobot-chromadb.service:39 is still a literal: --host 127.0.0.1. It is a static non-Jinja file, and :35-36 calls it a "static reference copy … kept in step". The 0.0.0.0 half of the criterion is met. The "from configuration" half is not met for this file.
    • Sweep: git grep -E -- '--host[ =]+0\.0\.0\.0' on main hits only the changelog and the guard's own synthetic fixtures.

    AC3 quotes:

    • 3a. All three files use loopback.
    • 3b. The decision is recorded inline at group_vars/all.yml:433-440: "Loopback by default (Option A): a multi-node install must set this explicitly and add a subnet-scoped firewall rule". It is echoed at redis/…service.j2:44-48.
    • 3c. The owner ruling reads: "Loopback-only is a dead end — clustering needs node-to-node reachability … The private cluster interface is the right default … Default the bind to the private cluster interface." The tree does not implement that default. Whether loopback-plus-explicit-override is acceptable for now is the owner's call.

    AC4 quotes:

    • 4a. The guard is repo_tests/chromadb_bind_not_hardcoded_15317_test.py, with _LAUNCH_SITE_PATTERNS = ("*.service", "*.service.j2", "*.sh") (:87).
    • 4b. It has a coverage floor (test_the_sweep_reaches_the_known_chroma_launch_sites, :326) and the assertion test_no_chroma_launch_site_hardcodes_a_wide_bind (:332).
    • 4c. It also has override, compose-port and host-networking checks (:374, :417, :430), each with a synthetic negative control (:442, :468, :477, :506, :523). CI runs repo_tests in .github/workflows/ci.yml:344 ("Run unit tests — backend, shared, tts-worker, repo_tests"). One limit: the guard targets wildcard binds. A literal non-wildcard like 2c's 127.0.0.1 passes it by design.

    AC5 quotes:

    Read vs inferred:

    • Read: the issue body and both comments; the lines quoted above from all three unit files, both defaults files, ai-stack.env.j2, the guard's test list and the docs section; and the repo_tests step in ci.yml.
    • Not run: the guard. I also did not read its full body; its coverage is taken from test names and the lines cited.
    • Inferred: that nothing else consumes chromadb_host as a bind. The grep was over autobot-slm-backend/ansible for chromadb_host: only.

    Undetermined:

    • Whether each deployment has a genuinely private segment. The owner ruling asked the fix to state this. I found no such statement in the tree, but searched only the files above.
    • What actually consumes the 0.0.0.0 dial default.

    Every criterion was reached; none was skipped.


    Generated by Claude Code

  6. mrveiss commented on Sep 28, 2026

    @mrveiss
    OwnerAuthor

    Verdict on the closure pass above: stays open on 2c. autobot-infrastructure/autobot-database/templates/autobot-chromadb.service:39 still carries a literal --host 127.0.0.1 in a static non-Jinja reference copy; the 0.0.0.0 half of the criterion is met, the from configuration half is not for that file. Either the reference copy is templated or the criterion states that a static reference copy is exempt — one line of work either way. Not closed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions