Parent
What to deliver
A versioned repetition policy justified by measured confidence stability rather than an arbitrary fixed count.
Blocked by
Procedure
- Produce at least one stable 200-request reference block.
- Repeatedly subsample 10, 20, 30, and 50 measured rows.
- Evaluate median estimation error, confidence-interval width, and winner-reversal rate for a matched comparison.
- Calibrate separate classes when needed: kernel microbenchmark, single-request engine, online p95/p99.
Required artifacts
- raw 200-request block;
- subsampling seeds and code;
- stability and reversal distributions;
repetition_policy.json or equivalent versioned artifact;
- validator/config update;
- report and verdict.
Acceptance criteria
Permitted claim after completion
The chosen repetition minima meet the calibrated stability target for the tested workload and environment class.
Prohibited claim
The selected count is not automatically sufficient for every model, tail quantile, or future software stack.
Parent
What to deliver
A versioned repetition policy justified by measured confidence stability rather than an arbitrary fixed count.
Blocked by
Procedure
Required artifacts
repetition_policy.jsonor equivalent versioned artifact;Acceptance criteria
Permitted claim after completion
The chosen repetition minima meet the calibrated stability target for the tested workload and environment class.
Prohibited claim
The selected count is not automatically sufficient for every model, tail quantile, or future software stack.