Skip to content

Release rehearsal: day-in-the-life simulation of 10 users (all roles) with container-resource + HTTP-traffic monitoring #243

Description

@mforce

This was generated by AI during triage.
Revised 2026-07-28 after a four-lens spec review (codex, architect agent, contrarian agent, owner direction): all five roles covered, weighted common-path pacing, hazard scenarios split from the clean baseline and documented, acceptance sharpened into thresholds. Review notes in the issue comments.

Goal

Rehearse a realistic production day before the first real users onboard. Produce a capacity/latency baseline for the prod-like single-container deployment, as input for release-rollout planning and the deployment-readiness audit (#244) — and a reusable, thresholded rehearsal script that can be rerun before every release.

Personas — all five sign-in kinds (src/Cluckwork.Domain/Accounts/Roles.cs)

10 users: 1 Owner, 1 Manager, 1 Sales, 3 Workers, 4 ReadOnly. Owner's stored/JWT role is Admin; Worker means no role row.

Count Sign-in Simulated day
1 Owner dashboard, reports, audit browse, one mid-day /export (the heaviest scans in the app, deliberately while writes are in flight), user-management touch, expenses/payments review
1 Manager daily-entry submit/review/adjust/void, flock ops, feed/water/inventory movements, stock corrections
1 Sales customers, draft orders, add lines, confirm, record payments — no production capture, no expenses (the only persona exercising SalesAccess)
3 Worker daily entries by grade, feed/water/mortality, culls; one of the three restricted to assigned flocks (permission narrowing exercised under load — its 403 probes against unassigned flocks are expected traffic)
4 ReadOnly dashboard, reports, stock/customer/sales browsing (read-heavy), plus a small weight of deep-links into /audit//export//users (the SPA has no route-level role gate — server-side 403s are real traffic)

Note: Lock is not a user action — it is the DailyEntryLockSweep background job (submitted entries older than 7 farm-local days, 30 s job loop). It participates via aged seed data (below), not via a persona.

Pacing — weighted common paths, not uniform random

  • Per-persona action-frequency table committed with the script (e.g. Worker: ~60% daily-entry save/read, 15% stock, 15% history, 10% misc; ReadOnly: ~40% dashboard, 30% reports, 20% stock, 10% history; Sales: ~35% orders, 25% customers, 20% payments, 20% stock). Weights seeded from the bottom-nav priority order (web/src/routes/nav.tsx).
  • Closed-model VUs with lognormal think time; staggered logins (no accidental stampede); the compressed day's wall-clock duration and compression factor are stated in the script README — they are the load-bearing parameters, and results are reported as request rates, not "10 users".
  • Run length ≥ 2× the 15-minute access-token lifetime so the refresh flow is actually exercised; refresh serialized per user (concurrent refresh burns the token family by design); per-VU cookie jar; X-Cluckwork-Auth CSRF header on refresh/logout.
  • Each screen modeled as its SPA call bundle (session bootstrap fires /me + /account; dashboard fans out), and static-asset traffic in scope — the same Kestrel serves both.
  • Idempotency keys: UUID per VU per logical operation, reused only on retry of that operation. Key scope is (AccountId, method+path, hash) — account-wide, not per-user — so cross-persona collisions silently swap cached responses; the script must make collisions impossible outside the deliberate replay probes.

Run design — three tagged phases, clean baseline kept clean

  1. Warmup (discarded from percentiles): JIT, EF model build, pool growth, migrations.
  2. Capacity baseline: the weighted persona day. Organic 409s recorded, nothing injected. ≥ 3 repetitions; report median with spread and a stated variance budget (±10% p95); reset between runs (docker compose down -v + reseed).
  3. Hazard passes (separate, each tagged, never mixed into baseline percentiles) — see matrix.

Hazard matrix — deliberately exercised, outcomes documented

Deliverable table in the findings doc: hazard → probe → expected outcome → observed → follow-up issue if divergent. Expected-count ranges are wired into the script's thresholds.

Hazard Probe Expected
Version-token race concurrent edit vs submit on one daily entry; concurrent AddItem on one order one wins, one 409; never silent lost update
FIFO stock contention two Sales confirm against one thin grade serialized; clean insufficient-stock failure; never negative stock
§4.6 currency-lock (#162) settings currency change racing the first financial row (needs a pristine second account — demo data is already locked) write wins or 422; never a re-priced total
Daily-entry natural key two Workers, same flock+date upsert semantics hold; no duplicate day
Idempotency replay same key re-sent after success; same key + different body cached response replayed; body not fingerprinted — documented as designed
Refresh reuse double-refresh inside the 10 s grace (OK) and outside it (family revoked → 401s → re-login) matches #169/#176 semantics
Auth rate limit (#143) deliberate login burst: request 11 in the window → 429 + Retry-After limiter fires; separate from baseline error budget
Withdrawal-restricted lot sale against a restricted lot blocked

Not simulated: the submit/lock race — the sweep only touches entries > 7 farm-local days old, unreachable in a compressed day; it is covered by aged seed data letting the sweep fire during the run, and the race class itself is pinned by the existing deterministic parallel-race integration tests (AGENTS.md). The matrix verifies each hazard class also has a deterministic integration test and records the gap if not.

Rate-limiter session model (decide up front, don't discover it mid-run)

Login is 10/900 s per client IP, no queue; refresh 60/900 s; plus per-account lockout (5 failures/15 min). One load-generator IP consumes the whole login window at exactly 10 users — the 2×/5× reruns cannot authenticate as specced. The script README states the chosen model: staggered pre-provisioned sessions for headroom runs (limits untouched — the run stays prod-config), with the deliberate 429 probe kept as its own hazard pass. Per-VU X-Forwarded-For is not used silently (the default compose trusts 172.16.0.0/12, so spoofing works — but then the run exercises a path no production client can take; if used, that trade-off is documented).

Seeded environment — dedicated Seed:Simulation profile (not Seed:Demo)

Seed:Demo cannot run this rehearsal: it creates neither the 10-user cast (base seeder = Owner + optional Worker only) nor meaningful volume (3 flocks, ~15 entries — reports and exports would measure empty tables).

  • Creates the full cast with exact roles + worker flock assignments; runtime-generated credentials, never committed secrets.
  • Parameterized history depth (e.g. 90/365 days): entries, lots with realistic FIFO depletion, orders/payments/expenses, populated audit table; actual row counts recorded in the findings header. Aged submitted entries let the lock sweep fire during the run.
  • Record states across the lifecycle: draft, submitted, locked, confirmed, partially paid, voidable.
  • A pristine second account for the currency-lock probe.
  • Deterministic reset recipe; farm timezone explicitly set (non-UTC), and at least one run straddles a farm-local or UTC midnight (Use farm-local dates for withdrawal restriction and allocation boundaries #35 boundary class).

Monitoring

Sizing runs — falsifiable, not laptop-shaped

The deploy compose sets no cpu/mem limits; an uncapped laptop run cannot back a sizing recommendation (GC sizes itself from visible memory; OOM-kill is the real failure mode in a cgroup). Capacity passes run at candidate instance shapes via compose override (e.g. 1 vCPU/1 GB and 2 vCPU/2 GB) and report pass/fail per shape; k6 pinned to disjoint CPUs or another machine; one verification pass through the traefik prod profile (TLS + XFF + limiter keying differ). The recommendation handed to #244 cites the capped runs only.

Deliverables

  • Parameterized k6 (or similar) script under tools/simulation/ — persona mix, weights, pacing, compression, user count, phase tags; README with exact commands, environment manifest fields (git SHA, image digest, host resources, config overrides, seed version + row counts, PRNG seed), and the chosen session model.
  • Seed:Simulation profile + deterministic reset recipe (as specced above).
  • Monitoring compose overlay under tools/simulation/.
  • Post-run DB invariant audit (script, committed): FIFO allocations ≤ lot quantities; stock = produced − allocated, never negative; ledger totals reconcile with entries; order totals = item sums; payments ≤ order totals; live refresh-token count as expected. This is the only oracle that catches the repo's own most-shipped bug class — a missing Version++ reduces 409s while silently losing writes.
  • Findings doc under docs/simulation/: resource headroom per candidate shape, top-N slowest endpoints, error/409/429 inventory, hazard matrix with observed outcomes, idempotency_records + audit_events growth measured across the run, run metadata header.
  • Sizing recommendation (from the capped runs) handed to Deployment readiness audit: production go-live scan before first real users (host undecided) #244.
  • Follow-up issues filed for anything divergent — known already: idempotency_records has no purge job (full response body stored per successful write, forever) — tracked separately; deploy compose has no restart: policy and no log rotation — handed to Deployment readiness audit: production go-live scan before first real users (host undecided) #244.

Acceptance — thresholds, not vibes

  • k6 thresholds encode the gates (script exits non-zero on breach): provisional p95 read < 500 ms, p95 write < 1 s, p99 < 2 s, export p95 < 5 s at seeded volume, sustained CPU < 70% at the candidate shape, no monotonic memory growth across the run; revisit with a documented decision if the baseline proves them wrong.
  • Baseline lane: zero unexpected 4xx/5xx (allow-list: organic 409 classes, the narrowing worker's 403s), zero missing-idempotency-key 400s, zero 429s (a 429 here means the session model is broken, and 429s are 4xx — "no 5xx" alone would pass a run where half the personas never logged in).
  • Hazard lane: exact expected status sets per matrix row; anything else fails the run.
  • Invariant audit passes after every run.
  • ≥ 3 clean baseline repetitions within the variance budget; 2×/5× characterization runs recorded.

Ongoing (this is a rehearsal, not a report)

  • CI smoke: 1 VU × 1 iteration per persona against the compose stack on workflow_dispatch + schedule (mirroring security-audit.yml), so the script's API contract cannot rot between releases; the full load run stays manual, rerun before each release tag. Each run's k6 summary JSON persisted for trend comparison.

Scope note

If this proves too large as one slice, the agreed split is: (1) script + seed profile + invariant audit + one clean capacity run; (2) monitoring overlay; (3) hazard passes; (4) capped sizing runs — each checkable off the phase epic separately.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions