You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Release rehearsal: day-in-the-life simulation of 10 users (all roles) with container-resource + HTTP-traffic monitoring #243
This was generated by AI during triage. Revised 2026-07-28 after a four-lens spec review (codex, architect agent, contrarian agent, owner direction): all five roles covered, weighted common-path pacing, hazard scenarios split from the clean baseline and documented, acceptance sharpened into thresholds. Review notes in the issue comments.
Goal
Rehearse a realistic production day before the first real users onboard. Produce a capacity/latency baseline for the prod-like single-container deployment, as input for release-rollout planning and the deployment-readiness audit (#244) — and a reusable, thresholded rehearsal script that can be rerun before every release.
Personas — all five sign-in kinds (src/Cluckwork.Domain/Accounts/Roles.cs)
10 users: 1 Owner, 1 Manager, 1 Sales, 3 Workers, 4 ReadOnly. Owner's stored/JWT role is Admin; Worker means no role row.
Count
Sign-in
Simulated day
1
Owner
dashboard, reports, audit browse, one mid-day /export (the heaviest scans in the app, deliberately while writes are in flight), user-management touch, expenses/payments review
customers, draft orders, add lines, confirm, record payments — no production capture, no expenses (the only persona exercising SalesAccess)
3
Worker
daily entries by grade, feed/water/mortality, culls; one of the three restricted to assigned flocks (permission narrowing exercised under load — its 403 probes against unassigned flocks are expected traffic)
4
ReadOnly
dashboard, reports, stock/customer/sales browsing (read-heavy), plus a small weight of deep-links into /audit//export//users (the SPA has no route-level role gate — server-side 403s are real traffic)
Note: Lock is not a user action — it is the DailyEntryLockSweep background job (submitted entries older than 7 farm-local days, 30 s job loop). It participates via aged seed data (below), not via a persona.
Pacing — weighted common paths, not uniform random
Per-persona action-frequency table committed with the script (e.g. Worker: ~60% daily-entry save/read, 15% stock, 15% history, 10% misc; ReadOnly: ~40% dashboard, 30% reports, 20% stock, 10% history; Sales: ~35% orders, 25% customers, 20% payments, 20% stock). Weights seeded from the bottom-nav priority order (web/src/routes/nav.tsx).
Closed-model VUs with lognormal think time; staggered logins (no accidental stampede); the compressed day's wall-clock duration and compression factor are stated in the script README — they are the load-bearing parameters, and results are reported as request rates, not "10 users".
Run length ≥ 2× the 15-minute access-token lifetime so the refresh flow is actually exercised; refresh serialized per user (concurrent refresh burns the token family by design); per-VU cookie jar; X-Cluckwork-Auth CSRF header on refresh/logout.
Each screen modeled as its SPA call bundle (session bootstrap fires /me + /account; dashboard fans out), and static-asset traffic in scope — the same Kestrel serves both.
Idempotency keys: UUID per VU per logical operation, reused only on retry of that operation. Key scope is (AccountId, method+path, hash) — account-wide, not per-user — so cross-persona collisions silently swap cached responses; the script must make collisions impossible outside the deliberate replay probes.
Run design — three tagged phases, clean baseline kept clean
Warmup (discarded from percentiles): JIT, EF model build, pool growth, migrations.
Capacity baseline: the weighted persona day. Organic 409s recorded, nothing injected. ≥ 3 repetitions; report median with spread and a stated variance budget (±10% p95); reset between runs (docker compose down -v + reseed).
Hazard passes (separate, each tagged, never mixed into baseline percentiles) — see matrix.
Deliverable table in the findings doc: hazard → probe → expected outcome → observed → follow-up issue if divergent. Expected-count ranges are wired into the script's thresholds.
Hazard
Probe
Expected
Version-token race
concurrent edit vs submit on one daily entry; concurrent AddItem on one order
one wins, one 409; never silent lost update
FIFO stock contention
two Sales confirm against one thin grade
serialized; clean insufficient-stock failure; never negative stock
deliberate login burst: request 11 in the window → 429 + Retry-After
limiter fires; separate from baseline error budget
Withdrawal-restricted lot
sale against a restricted lot
blocked
Not simulated: the submit/lock race — the sweep only touches entries > 7 farm-local days old, unreachable in a compressed day; it is covered by aged seed data letting the sweep fire during the run, and the race class itself is pinned by the existing deterministic parallel-race integration tests (AGENTS.md). The matrix verifies each hazard class also has a deterministic integration test and records the gap if not.
Rate-limiter session model (decide up front, don't discover it mid-run)
Login is 10/900 s per client IP, no queue; refresh 60/900 s; plus per-account lockout (5 failures/15 min). One load-generator IP consumes the whole login window at exactly 10 users — the 2×/5× reruns cannot authenticate as specced. The script README states the chosen model: staggered pre-provisioned sessions for headroom runs (limits untouched — the run stays prod-config), with the deliberate 429 probe kept as its own hazard pass. Per-VU X-Forwarded-For is not used silently (the default compose trusts 172.16.0.0/12, so spoofing works — but then the run exercises a path no production client can take; if used, that trade-off is documented).
Seed:Demo cannot run this rehearsal: it creates neither the 10-user cast (base seeder = Owner + optional Worker only) nor meaningful volume (3 flocks, ~15 entries — reports and exports would measure empty tables).
Creates the full cast with exact roles + worker flock assignments; runtime-generated credentials, never committed secrets.
Parameterized history depth (e.g. 90/365 days): entries, lots with realistic FIFO depletion, orders/payments/expenses, populated audit table; actual row counts recorded in the findings header. Aged submitted entries let the lock sweep fire during the run.
Record states across the lifecycle: draft, submitted, locked, confirmed, partially paid, voidable.
A pristine second account for the currency-lock probe.
Container resources: sampled docker stats (CPU, memory, net + block I/O, DB volume growth) for app and Postgres.
HTTP: k6 trends tagged by persona/flow/endpoint/hazard; p50/p95/p99; status mix. OTLP meters (Observability: metrics — runtime, HTTP, and DB meters over OTLP #215) with export interval dropped to ~10 s for the run; filter /health out of request-duration reads. OTLP traces too (ASP.NET + EF spans already emitted).
Postgres side: pg_stat_statements before/after, pg_stat_activity/lock-wait + pg_blocking_pids sampling. Npgsql Max Pool Size (100) sits exactly at stock max_connections (100) — pool saturation is a plausible first failure and must be visible.
The deploy compose sets no cpu/mem limits; an uncapped laptop run cannot back a sizing recommendation (GC sizes itself from visible memory; OOM-kill is the real failure mode in a cgroup). Capacity passes run at candidate instance shapes via compose override (e.g. 1 vCPU/1 GB and 2 vCPU/2 GB) and report pass/fail per shape; k6 pinned to disjoint CPUs or another machine; one verification pass through the traefik prod profile (TLS + XFF + limiter keying differ). The recommendation handed to #244 cites the capped runs only.
Deliverables
Parameterized k6 (or similar) script under tools/simulation/ — persona mix, weights, pacing, compression, user count, phase tags; README with exact commands, environment manifest fields (git SHA, image digest, host resources, config overrides, seed version + row counts, PRNG seed), and the chosen session model.
Monitoring compose overlay under tools/simulation/.
Post-run DB invariant audit (script, committed): FIFO allocations ≤ lot quantities; stock = produced − allocated, never negative; ledger totals reconcile with entries; order totals = item sums; payments ≤ order totals; live refresh-token count as expected. This is the only oracle that catches the repo's own most-shipped bug class — a missing Version++reduces 409s while silently losing writes.
Findings doc under docs/simulation/: resource headroom per candidate shape, top-N slowest endpoints, error/409/429 inventory, hazard matrix with observed outcomes, idempotency_records + audit_events growth measured across the run, run metadata header.
k6 thresholds encode the gates (script exits non-zero on breach): provisional p95 read < 500 ms, p95 write < 1 s, p99 < 2 s, export p95 < 5 s at seeded volume, sustained CPU < 70% at the candidate shape, no monotonic memory growth across the run; revisit with a documented decision if the baseline proves them wrong.
Baseline lane: zero unexpected 4xx/5xx (allow-list: organic 409 classes, the narrowing worker's 403s), zero missing-idempotency-key 400s, zero 429s (a 429 here means the session model is broken, and 429s are 4xx — "no 5xx" alone would pass a run where half the personas never logged in).
Hazard lane: exact expected status sets per matrix row; anything else fails the run.
Invariant audit passes after every run.
≥ 3 clean baseline repetitions within the variance budget; 2×/5× characterization runs recorded.
Ongoing (this is a rehearsal, not a report)
CI smoke: 1 VU × 1 iteration per persona against the compose stack on workflow_dispatch + schedule (mirroring security-audit.yml), so the script's API contract cannot rot between releases; the full load run stays manual, rerun before each release tag. Each run's k6 summary JSON persisted for trend comparison.
Scope note
If this proves too large as one slice, the agreed split is: (1) script + seed profile + invariant audit + one clean capacity run; (2) monitoring overlay; (3) hazard passes; (4) capped sizing runs — each checkable off the phase epic separately.
Goal
Rehearse a realistic production day before the first real users onboard. Produce a capacity/latency baseline for the prod-like single-container deployment, as input for release-rollout planning and the deployment-readiness audit (#244) — and a reusable, thresholded rehearsal script that can be rerun before every release.
Personas — all five sign-in kinds (
src/Cluckwork.Domain/Accounts/Roles.cs)10 users: 1 Owner, 1 Manager, 1 Sales, 3 Workers, 4 ReadOnly. Owner's stored/JWT role is
Admin; Worker means no role row./export(the heaviest scans in the app, deliberately while writes are in flight), user-management touch, expenses/payments reviewSalesAccess)/audit//export//users(the SPA has no route-level role gate — server-side 403s are real traffic)Note: Lock is not a user action — it is the
DailyEntryLockSweepbackground job (submitted entries older than 7 farm-local days, 30 s job loop). It participates via aged seed data (below), not via a persona.Pacing — weighted common paths, not uniform random
web/src/routes/nav.tsx).X-Cluckwork-AuthCSRF header on refresh/logout./me+/account; dashboard fans out), and static-asset traffic in scope — the same Kestrel serves both.(AccountId, method+path, hash)— account-wide, not per-user — so cross-persona collisions silently swap cached responses; the script must make collisions impossible outside the deliberate replay probes.Run design — three tagged phases, clean baseline kept clean
docker compose down -v+ reseed).Hazard matrix — deliberately exercised, outcomes documented
Deliverable table in the findings doc: hazard → probe → expected outcome → observed → follow-up issue if divergent. Expected-count ranges are wired into the script's thresholds.
AddItemon one orderRetry-AfterNot simulated: the submit/lock race — the sweep only touches entries > 7 farm-local days old, unreachable in a compressed day; it is covered by aged seed data letting the sweep fire during the run, and the race class itself is pinned by the existing deterministic parallel-race integration tests (AGENTS.md). The matrix verifies each hazard class also has a deterministic integration test and records the gap if not.
Rate-limiter session model (decide up front, don't discover it mid-run)
Login is 10/900 s per client IP, no queue; refresh 60/900 s; plus per-account lockout (5 failures/15 min). One load-generator IP consumes the whole login window at exactly 10 users — the 2×/5× reruns cannot authenticate as specced. The script README states the chosen model: staggered pre-provisioned sessions for headroom runs (limits untouched — the run stays prod-config), with the deliberate 429 probe kept as its own hazard pass. Per-VU
X-Forwarded-Foris not used silently (the default compose trusts172.16.0.0/12, so spoofing works — but then the run exercises a path no production client can take; if used, that trade-off is documented).Seeded environment — dedicated
Seed:Simulationprofile (notSeed:Demo)Seed:Democannot run this rehearsal: it creates neither the 10-user cast (base seeder = Owner + optional Worker only) nor meaningful volume (3 flocks, ~15 entries — reports and exports would measure empty tables).Monitoring
docker stats(CPU, memory, net + block I/O, DB volume growth) for app and Postgres./healthout of request-duration reads. OTLP traces too (ASP.NET + EF spans already emitted).pg_stat_statementsbefore/after,pg_stat_activity/lock-wait +pg_blocking_pidssampling. Npgsql Max Pool Size (100) sits exactly at stockmax_connections(100) — pool saturation is a plausible first failure and must be visible.tools/simulation/— and note for Deployment readiness audit: production go-live scan before first real users (host undecided) #244: prod-as-deployed exports nothing (Otlp:Endpointunset ⇒ disabled); whether prod ships with an OTLP destination is an explicit Deployment readiness audit: production go-live scan before first real users (host undecided) #244 decision, not this issue's side effect.Sizing runs — falsifiable, not laptop-shaped
The deploy compose sets no cpu/mem limits; an uncapped laptop run cannot back a sizing recommendation (GC sizes itself from visible memory; OOM-kill is the real failure mode in a cgroup). Capacity passes run at candidate instance shapes via compose override (e.g. 1 vCPU/1 GB and 2 vCPU/2 GB) and report pass/fail per shape; k6 pinned to disjoint CPUs or another machine; one verification pass through the traefik
prodprofile (TLS + XFF + limiter keying differ). The recommendation handed to #244 cites the capped runs only.Deliverables
tools/simulation/— persona mix, weights, pacing, compression, user count, phase tags; README with exact commands, environment manifest fields (git SHA, image digest, host resources, config overrides, seed version + row counts, PRNG seed), and the chosen session model.Seed:Simulationprofile + deterministic reset recipe (as specced above).tools/simulation/.Version++reduces 409s while silently losing writes.docs/simulation/: resource headroom per candidate shape, top-N slowest endpoints, error/409/429 inventory, hazard matrix with observed outcomes,idempotency_records+audit_eventsgrowth measured across the run, run metadata header.idempotency_recordshas no purge job (full response body stored per successful write, forever) — tracked separately; deploy compose has norestart:policy and no log rotation — handed to Deployment readiness audit: production go-live scan before first real users (host undecided) #244.Acceptance — thresholds, not vibes
thresholdsencode the gates (script exits non-zero on breach): provisional p95 read < 500 ms, p95 write < 1 s, p99 < 2 s, export p95 < 5 s at seeded volume, sustained CPU < 70% at the candidate shape, no monotonic memory growth across the run; revisit with a documented decision if the baseline proves them wrong.Ongoing (this is a rehearsal, not a report)
workflow_dispatch+ schedule (mirroringsecurity-audit.yml), so the script's API contract cannot rot between releases; the full load run stays manual, rerun before each release tag. Each run's k6 summary JSON persisted for trend comparison.Scope note
If this proves too large as one slice, the agreed split is: (1) script + seed profile + invariant audit + one clean capacity run; (2) monitoring overlay; (3) hazard passes; (4) capped sizing runs — each checkable off the phase epic separately.