Repository navigation
idempotency_records grows unbounded: full response body stored per successful write, no purge job #259
Description
Activity
Measured, not estimated — #243 harness on merged
main(5df5cd7)The open decision was whether the purge sweep needs a
CreatedAtindex. Here are numbers instead of arithmetic.Instrument note (this part matters)
seed --profile simulationwrites zeroidempotency_recordsrows — it invokes handlers directly through DI, with noHttpClientand noIdempotency-Key, andIdempotencyMiddlewareis HTTP-layer. Confirmed empirically: a full 90-day fixture seed left the table at 0 rows. Measuring with the seeder would have produced a confident "no index needed" answer to a different question. All numbers below come from the k6 harness driving real HTTP writes.The run
10 VUs (full cast: Owner/Manager/Sales/3×Worker/4×ReadOnly), 10-minute capacity phase, prod-config container.
Capacity requests 8,480 (13.33 req/s) checks rate / unexpected statuses 1.0 / 0 Per-cast-user coverage all 10, exactly once each Growth
Rows written 0 → 362 in 596 s Write rate 0.607 rows/s (36.4/min) Share of traffic 362 / 8,480 = 4.3% of requests are idempotent writes Total relation size 376 kB (heap 200 kB + indexes 144 kB) Bytes per row 1,063 B The premise needs correcting
This issue's framing is "full response body stored per successful write", implying bodies dominate. They don't:
ResponseBodyAverage 68 bytes Max 99 bytes Total across 362 rows 24 kB — 6.4% of the table The table is dominated by row overhead, the three hash columns, and the two existing indexes — not by cached bodies. This app's write endpoints return small bodies (created ids and status), so the "unbounded response-body storage" concern is much weaker than filed. The growth concern is real; the reason stated for it is not.
Does the sweep need an index?
Synthesized the 48h steady state implied by the measured rate — 105,000 rows / 33 MB — with the real column shape, and ran the proposed batched delete both ways. Worst case deliberately: no row is old enough yet, so the
LIMIT 1000never short-circuits and the scan runs to completion.Execution time No index (today's schema) 16.9 ms — full Seq Scan, 105k rowsCreatedAtbtree index0.062 ms — Index ScanIndex cost 2,328 kB + write amplification on every idempotent write Recommendation: code-only sweep, no index, and do not touch
InitialCreateThe index is 273× faster in relative terms and irrelevant in absolute ones. 17 ms once per sweep interval on a background worker is free; the index would tax every write on the hot idempotency insert path to buy it back. That trade is backwards.
Two honesty caveats on the extrapolation:
- 0.607 rows/s is a synthetic ceiling, not a farm. It is 10 users acting continuously with ~1.5 s think times, 24/7. Sustained naively that is ~52k rows/day → ~111 MB at a 48h window. A real farm writes in bursts during working hours, so the true steady state is orders of magnitude smaller. The unindexed scan is comfortable even at the ceiling, which is the point.
- Sim-stack deviations apply (per
tools/simulation/README.md): plaintext local Postgres, uncapped host resources, raised rate limits. Absolute latency here is not production-equivalent — but the scan-vs-index comparison is a same-machine A/B, so the 273× ratio and the 17 ms order of magnitude hold.
Blocker cleared along the way
The harness could not boot merged
mainat all — four breakages (#319AllowedHosts, #261/#262 TLS floor, abootstrap.shheredoc executing a comment, and a k6 timezone bug that failed 12.4% of requests during the hours UTC and the farm's date disagree). Fixed in a separate PR with a CI smoke so this cannot rot silently again.- added a commit that references this issue
on Aug 5, 2026
Found by the #243 spec review (contrarian + architect lenses, verified against source).
The gap
IdempotencyMiddlewarestores the full response body of every successful (2xx) write inidempotency_records, keyed by(AccountId, method+path, SHA256(key)). Nothing ever deletes these rows: the Jobs folder contains onlyDailyEntryLockSweep,DurableJobWorker", and its heartbeat — no purge readsCreatedAt.audit_events` grows the same way, but that is arguably a feature (audit retention is a product decision); cached idempotency responses are not — their usefulness ends when a client retry is no longer plausible (hours, not months).Every production write grows the table monotonically, with response bodies inline. At farm scale this is slow-burn, but it is unbounded, and the #243 rehearsal will quantify the growth rate (its findings doc records before/after table sizes).
Suggested shape
DurableJobWorkerloop deleting records older than a retention window (config, default e.g. 48 h — comfortably beyond any client retry horizon, aligned with how idempotent replays are actually used by the SPA).audit_eventsretention is in scope here or its own product decision (lean: own decision, not this issue).Refs
src/Cluckwork.Api/Middleware/IdempotencyMiddleware.cs(storage, key scope)src/Cluckwork.Infrastructure/Persistence/IdempotencyRecord.cs(unique index, CreatedAt)src/Cluckwork.Infrastructure/Jobs/(the worker loop to hang the sweep on)