Skip to content

epic(settings): multi-tenant settings directory and tenant-scoped runtime #583

Description

@taitelee

Area: settings / api / streaming — multi-tenant preparation epic · decided with @EricAndrechek, 2026-09-10 → 2026-09-15

WaveHouse runs one settings snapshot per process. WaveHouse Cloud needs one process to serve many tenants, each with its own roles, policies, pipes, and config, delivered by the cloud fan-out as files on disk exactly like standalone. This epic makes the settings directory, the in-memory snapshots, and every request path tenant-aware without changing anything for a standalone operator. Split into the stories below as they are promoted to Ready (the #194 pattern).

Decisions

  • Two directory shapes, no hybrid. Flat: the four files (config.json, policies.json, pipes.json, roles.json) sit in the root and are tenant 0. Nested: the root holds only subfolders, each folder name is a tenant id, each folder holds the four files. Loose files and subfolders in the same root is an error at boot and at reload. Switching shapes is stop, restructure, start.
  • Tenant id is a validated string. One rule: safe as a folder name and safe as a NATS subject token (letters, digits, _, -; no spaces, dots, slashes, or wildcards; capped length). The folder name is the id. 0 is the reserved default. Not an integer: 19-digit ids already round when treated as numbers (internal/auth/auth.go).
  • Requests resolve to a tenant before auth. The id comes from a request header (name is a constant, TBD); absent means 0. In flat mode tenant 0 always exists. In nested mode a 0 folder is allowed but advised against; the control plane won't create one. Putting the header on every request is the client's or proxy's job.
  • Settings owns the tenant map. The registry maps tenant id → *settings.Store. The middleware resolves the header against it and 404s on a miss; what goes into the request context is the resolved store (the long-lived handle whose getters read the latest adopted document), not the bare id. Handlers extract it once and pass it down as an explicit parameter; nothing below a handler reads context. Async paths (ingest worker, sweeper, DLQ, stream hub) carry the tenant in the MQ subject.
  • Invalid tenant folder: fail closed per tenant, per pod. At boot and on reload alike, a folder that fails Validate drops that tenant on that pod (its requests 503 with the findings); every other tenant on that pod, and every other pod, keeps serving. No previous-snapshot fallback in nested mode. The control plane runs the same settings.Validate (pure, no deps) before fan-out, so an invalid folder on disk is a bug, not a normal event. Flat mode keeps today's behavior: refuse to boot, keep the previous snapshot on a rejected reload. No status route for now.
  • Reload triggers. Flat mode keeps fsnotify + SIGHUP + POST /v1/ops/settings/reload. Nested mode does not use fsnotify: the control plane writes the folder and calls the reload route. SIGHUP works in both modes and always reloads the whole tree. The route takes an optional tenant id: absent reloads the whole tree, present looks the tenant up and reloads that folder. Ops routes stay tenant-exempt (/livez, /readyz, metrics, version, /v1/ops/*).
  • Ops routes take the operator key only in nested mode. Today the /v1/ops/* gate admits the operator key or a token whose role is the policy's admin role. In nested mode there is no single policy to read an admin role from, and no tenant's admin may act on another tenant, so the gate admits the operator key alone and a token admin gets 403: tenant admins never touch the ops routes. The operator key is boot-level and reaches every ops route for every tenant, whether the call names one (?tenant=) or none (the whole tree). Flat mode is unchanged — the gate still admits tenant 0's admin role — so nothing changes for a standalone operator. Decided for now (2026-09-21); a per-tenant admin surface would be a later story.
  • Scope stays as is. scope is a within-table partition (an org inside a shared table); tenant selects the settings and the tables. Different axes. Scope remains inert and tenant goes in as a separate leading token: cache namespaces tenant.table.version.scope, stream topic tenant.table[.scope], ingest subjects ingest.<tenant>.<table> / dlq.<tenant>.<table>.
  • Every folder carries a full config.json. ClickHouse, auth, dedupe, DLQ, query, schema, stream, MQ, CORS are all per tenant, so the runtime resources behind them become per tenant too (stories 6–9).
  • Cache is one global pool. One ristretto instance, tenant-prefixed keys. Heavier tenants hold more of it. Redis is the next layer, later.
  • Settled with Eric on 2026-09-21, landed in feat(settings): nested settings directory with a per-tenant registry #598. A rejected tenant's 503 is generic. The reload response body is unchanged and a 422 can mean adopted in part; a failure about the directory itself rejects the reload whole. One shared keepalive wheel at the shortest interval among served tenants, until feat(stream): give each tenant its own keepalive wheel #738 gave each tenant its own (feat(stream): honor each tenant's keepalive settings on the wheel #597). A badly named folder is skipped, never fatal. A whole-tree reload mirrors the folders (new adopted, missing dropped). Dependencies.PolicySource stays and is ignored when the registry is nested. /v1/ops/* admits the operator key alone in nested mode; boot warns when a nested directory has no operator key.
  • Landed in feat(mq): lead every subject with the tenant #609 (story 5a), decided 2026-09-23; Eric's to revisit. The tenant token goes in verbatim, since the id grammar already makes it one subject token. A subject without one, written before the upgrade, reads as tenant 0, so an upgrade needs no drain and rows parked before it count as tenant 0's. Until 5b, the sweeper's longest gap window among the tenants served is the keepalive wheel's rule: the one value that satisfies every tenant. Say acme keeps 15 minutes and globex 60: acme's history stays 60 minutes too, counted against the one shared mq.max_bytes_gb.
  • Landed in feat(clickhouse): tls, headers, pool sizes and a connection ceiling #603 and feat(clickhouse): one pool per tuple and one schema registry per tenant #610 (story 6), decided 2026-09-21 → 2026-09-23; Eric's to revisit. One native ClickHouse pool per distinct address, database, user, password and tls tuple among the tenants served, shared by the tenants naming it and sized to their largest ask. clickhouse.max_total_conns caps the open pools together: boot refuses above it, and a reload that would exceed it is refused and logged, the tenant keeping its previous pool. A tenant on no pool answers 503 with Retry-After: 30, and a table lookup before its tenant's first discovery answers 503 with Retry-After: 5 where it was a 404. Over a nested directory /livez is 503 until some tenant completes a first discovery and 200 for the rest of the process lifetime, and /readyz is ready when any open pool answers. An insert invalidates cached results for every tenant on the same address and database, and a tenant readmitted after a rejection or removal, or moved to another address or database, has its cached structured-query results dropped at once. Per-tenant ClickHouse passwords wait for auth(config): where boot secrets live and how they rotate (jwt_secret, operator_key, clickhouse.password) #529.
  • Landed in feat(app): end a removed tenant's streams, one Pebble for every tenant #611 (story 3), decided with Eric 2026-09-24. A tenant that stops being served, removed or rejected, has its open streams ended; the reconnect gets 404 or 503, and the SDK stops on the first and retries the second, resuming from Last-Event-ID. /v1/health keeps resolving a tenant: every tenant route answers an unknown tenant before authentication anyway, and a tenant's existence is not a secret (Cloud's proxy is meant to turn away unknown subdomains). A removed or rejected tenant's queued rows go to the dead-letter queue rather than staying unacked, where they would hold the shared stream's ack floor, and Eric expects that to hold with per-tenant streams too. Dedupe is one Pebble instance for every tenant, keys led by the tenant, the layout decided inside the embedded implementation.
  • Landed in feat(mq): give every tenant a queue of its own #612 (story 5b), talked through with Eric 2026-09-24. Every tenant has a stream pair of its own inside the embedded implementation: INGEST_<tenant> at the tenant's own mq.max_bytes_gb and DLQ_<tenant> at a tenth of it, opened with the tenant's first budget and kept on removal, so a removed or rejected tenant's queued rows are still delivered and parked on its own dead-letter queue. The sweeper purges each tenant at its own gap window: a rejected tenant keeps the window its folder last had (all of its history when rejected since boot), and a removed tenant keeps no acknowledged history. GET /v1/ops/dlq/stats reads one tenant, tenant 0 when ?tenant= is absent, looked up in the MQ rather than the settings. A budget is a cap and never a disk reservation; what the budgets add up to against the disk is mq: validate WH_MQ_MAX_BYTES_GB upper bound to prevent disk over-reservation #138's. A reload that shrinks a budget never caps a dead-letter stream below what it holds (mq(dlq): shrinking mq.max_bytes_gb silently deletes the oldest dead letters — investigate how the reload should treat a non-empty DLQ #532, interim). Boot deletes the old shared WAVEHOUSE and WAVEHOUSE_DLQ streams. One tenant's failed queue join costs that tenant alone (fix(mq): keep one tenant's failed queue join from ending ingest for all #680).
  • Deferred by decision. Boot secrets (auth(config): where boot secrets live and how they rotate (jwt_secret, operator_key, clickhouse.password) #529). Dedupe remote backend (Scylla/Dynamo), dedupe admin routes, purge and rename of tenants, the per-tenant status route, a ClickHouse socket ceiling beyond a process-wide cap, and a scheduler or worker model for the background loops — discovery, JWKS refresh, the sweeper, batch inserts (per-tenant loops with jitter instead for now; Eric, 2026-09-24, wants them reworked together later; a library like gocron v2 if a cap is ever needed).

Stories

Order

Out of scope

Related: #508 (settings reload), #530 (boot fails loudly), #140, #138, #532, #214, #235, #262, #361.

Activity

  1. added
    enhancementNew feature or request
    area/apiHTTP handlers, routing, middleware
    area/streamingSSE / live-query delivery path (/v1/stream)
    area/configConfig file, config knobs, hot-reload
    on Sep 10, 2026
  2. moved this from Backlog to In progress in WaveHouse Task Boardon Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/apiHTTP handlers, routing, middlewarearea/configConfig file, config knobs, hot-reloadarea/streamingSSE / live-query delivery path (/v1/stream)enhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions