Problem
This resolves #73 (perf(cache): query cache returns stale data after writes).
Following our architecture review, we explicitly rejected the idea of writing a mini query-engine to merge NATS WAL buffers with live queries. We are adopting an "acceptable delay" paradigm for structured reads, but enforcing strict Read-Your-Writes guarantees.
Currently, the API cache relies on a static TTL. If Bento flushes a write batch at second 1, the user still sees stale data until the TTL expires, leading to test collisions and administrative confusion. Additionally, because cache keys are SHA256 hashes of SQL strings, prefix-based invalidation was non-viable.
Proposed Solution
This PR implements the two-pronged approach established in the architecture review (Option B + Option C):
-
Option C: Raw SQL Cache Bypass:
- Caching has been completely stripped from
QueryHandler.Handle (/v1/query).
- Admin raw-SQL queries (used heavily in E2E tests and debugging) will now always hit ClickHouse directly, guaranteeing immediate consistency without cache thrash.
-
Option B: Per-Table Tagging and Invalidation:
- Modified the
cache.Cache interface to accept tags []string during Set(), and replaced prefix invalidation with InvalidateByTags(ctx, tags).
LocalCache now maintains a secondary tagsMap (mapping a tag to its generated SHA256 cache keys).
- In
pipes.go, structured queries use a naive regex parser to extract FROM and JOIN table names, attaching them as tags to the cache entry.
- In
bento.go, the ingest worker now accepts the TieredCache. Upon a successful 200 OK from ClickHouse, Bento immediately drops all cache entries tagged with the inserted table_name.
Acceptance Criteria
Additional Context
- The naive SQL parser is a stop-gap for v1. It may over-match (false positives -> safe cache invalidation) but aims to avoid false negatives.
- If high-write tables experience cache thrash, we can tune this later with time-bucketed aggregation strategies.
Problem
This resolves #73 (
perf(cache): query cache returns stale data after writes).Following our architecture review, we explicitly rejected the idea of writing a mini query-engine to merge NATS WAL buffers with live queries. We are adopting an "acceptable delay" paradigm for structured reads, but enforcing strict Read-Your-Writes guarantees.
Currently, the API cache relies on a static TTL. If Bento flushes a write batch at second 1, the user still sees stale data until the TTL expires, leading to test collisions and administrative confusion. Additionally, because cache keys are SHA256 hashes of SQL strings, prefix-based invalidation was non-viable.
Proposed Solution
This PR implements the two-pronged approach established in the architecture review (Option B + Option C):
Option C: Raw SQL Cache Bypass:
QueryHandler.Handle(/v1/query).Option B: Per-Table Tagging and Invalidation:
cache.Cacheinterface to accepttags []stringduringSet(), and replaced prefix invalidation withInvalidateByTags(ctx, tags).LocalCachenow maintains a secondarytagsMap(mapping a tag to its generated SHA256 cache keys).pipes.go, structured queries use a naive regex parser to extractFROMandJOINtable names, attaching them as tags to the cache entry.bento.go, the ingest worker now accepts theTieredCache. Upon a successful200 OKfrom ClickHouse, Bento immediately drops all cache entries tagged with the insertedtable_name.Acceptance Criteria
/v1/query) bypass the cache entirely (returningX-Cache: BYPASS).Cacheinterface andLocalCachesupport tagging and tag-based invalidation.clickhouseOutputinvalidates cache tags synchronously only upon successful ClickHouse insert.Additional Context