Skip to content

E03: calibrate repetition counts and winner-reversal risk #8

Description

@hanklin9188

Parent

What to deliver

A versioned repetition policy justified by measured confidence stability rather than an arbitrary fixed count.

Blocked by

Procedure

  • Produce at least one stable 200-request reference block.
  • Repeatedly subsample 10, 20, 30, and 50 measured rows.
  • Evaluate median estimation error, confidence-interval width, and winner-reversal rate for a matched comparison.
  • Calibrate separate classes when needed: kernel microbenchmark, single-request engine, online p95/p99.

Required artifacts

  • raw 200-request block;
  • subsampling seeds and code;
  • stability and reversal distributions;
  • repetition_policy.json or equivalent versioned artifact;
  • validator/config update;
  • report and verdict.

Acceptance criteria

  • The selected engine minimum keeps winner-reversal risk below 5% for the calibrated comparison class.
  • Kernel and online-tail minima are separately justified or explicitly inherited with limitations.
  • The validator consumes the versioned policy rather than retaining unexplained hard-coded counts.
  • Fewer-than-required fixtures cannot receive PASS.
  • Increasing repetitions is not used to average over systematic thermal drift.

Permitted claim after completion

The chosen repetition minima meet the calibrated stability target for the tested workload and environment class.

Prohibited claim

The selected count is not automatically sufficient for every model, tail quantile, or future software stack.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions