Skip to content

benchd: seal the timed-token tolerance figures only coarsened - #5

Open
GumbiiDigital wants to merge 1 commit into
Layr-Labs:mainfrom
GumbiiDigital:upstream/seal-token-counts
Open

GumbiiDigital wants to merge 1 commit into
Layr-Labs:mainfrom
GumbiiDigital:upstream/seal-token-counts

Conversation

@GumbiiDigital

Copy link
Copy Markdown
Collaborator

Summary

The timed-token tolerance path added in the 2026-09-28 snapshot seals engine-chosen values verbatim:

  • token_mismatch_count, token_mismatch_first_step, token_mismatch_near_tie_count (and the second-choice count and max relative gap), in the flat metrics and in every paired_legs row, including refused runs;
  • the TIMED-DIVERGENCE-NOT-A-NEAR-TIE text, which interpolates the first diverging position and the reference gap at full precision, and the TIMED-DIVERGENCE-OVER-TOLERANCE text, which interpolates the candidate's mismatch count.

The sealed record goes back to the participant, whose engine has seen the hidden prompt and chooses which tokens it commits. Exact values therefore form a small read-back channel (several bits per pair).

Change

  • Sealed record: the counts seal as a power-of-two bucket (0 stays 0), the max second-choice gap at 2 decimals, and token_mismatch_first_step not at all (absent on the flat metrics, null on a pair row, so the row's key set is unchanged). The judge still decides on the exact in-memory values.
  • Refusal texts: NOT-A-NEAR-TIE names no position and rounds the gap to 2 decimals. OVER-TOLERANCE states the allowance benchd computes from the fixture instead of the candidate's count. The tolerance refusal seals first_failing_step as null.
  • Logging: exact values go to stderr for operators.

2 files: crates/benchd/src/official.rs, crates/benchd/src/score.rs.

Tests

  • New the_sealed_record_holds_no_exact_token_counts: count 13, first step 45 and gap 0.123456 seal as 16, no step, and 0.12, on the flat metrics and a pair row.
  • Updated the key-pin bytes, the additive key set and the two refusal-text assertions.
  • cargo fmt --check, cargo clippy --workspace --all-targets -D warnings and all benchd tests pass (655, 1 ignored).
  • A few bench-runner process-spawn tests (not touched here) fail intermittently under parallel runs and pass with --test-threads=1.

🤖 Generated with Claude Code

The timed-token tolerance path sealed the candidate's token_mismatch_*
figures verbatim, and its refusal texts interpolated the candidate's
mismatch count, the first diverging position and the reference gap at
full precision. The sealed record goes back to the participant whose
engine has seen the hidden prompt, and the engine chooses which tokens
it commits, so those exact values form a small read-back channel.

- Sealed record (flat metrics and every paired_legs row, refused runs
  included): the mismatch, near-tie and second-choice counts seal as a
  power-of-two bucket (0 stays 0), the max second-choice gap at 2
  decimals, and token_mismatch_first_step not at all (absent on the flat
  metrics, null on a pair row). In memory they stay exact for the judge.
- TIMED-DIVERGENCE-NOT-A-NEAR-TIE names no position and rounds the gap
  to 2 decimals; TIMED-DIVERGENCE-OVER-TOLERANCE states the allowance
  benchd computed from the fixture instead of the candidate's count. The
  tolerance refusal seals first_failing_step as null.
- benchd logs the exact values to stderr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant