Skip to content

eval: the judge scores identical answers differently, which is what makes the passing set unstable #324

Description

@antejavor

What

The GEval judge gives different scores to the same answer text on comparable questions, and the disagreement straddles the pass threshold.

From a single run's abstention questions — same rubric, same judge, same model, one run:

score verdict answer
0.7 PASS '0'
0.7 PASS '0'
0.7 PASS '0'
0.7 PASS 'not in memory'
0.7 PASS 'not in memory'
0.2 fail 'not in memory'
0.2 fail 'not in memory'
0.0 fail 'You lead a team of 4 engineers…' (genuine fabrication)

The exact string not in memory scored 0.2, 0.2, 0.7, 0.7. Half passed, half failed, on questions where the expected output is the same shape of "this is not in memory".

The 0.0 is correct and shows the judge is not simply broken — it distinguishes a fabricated answer from a refusal. It just doesn't score refusals consistently.

Why this matters more than it looks

This is the missing explanation for the calibration result. Three repeat runs on an identical, frozen graph gave:

3 runs: 15%, 10%, 15%   noise floor +/-5pp
stable passes: 0 of 6 questions that passed at least once

Zero questions passed in all three runs, and the first two runs' passing sets were entirely disjoint. That was previously attributed to retrieval nondeterminism. At least part of it is the judge: with the graph fixed and retrieval producing the same answer, the score still moves across the gate.

The consequence is that a pass/fail flip does not imply a behaviour change, so:

  • coverage deltas below the floor are judge noise, not retrieval noise
  • the efficiency median — taken over whichever questions passed — is resampled by the judge, not by retrieval
  • compare() can attribute a regression to a schema change that never happened

Notes on the threshold

Scores cluster heavily: across four runs, 42/80 rubric scores are exactly 0.0 and 9 are exactly 1.0, with a thin spread between. Six land exactly on the 0.7 gate — i.e. the passes are at the threshold, not above it. A metric whose passes sit exactly on the boundary will flip on the smallest judge variation.

Directions worth considering

  • Score each question more than once and take a median, at the cost of N× judge calls. Expensive but it directly attacks the variance.
  • Move the gate off the cluster point — if passes land exactly on 0.7, a 0.7 threshold is the worst possible choice.
  • Pin sampling if the judge exposes temperature/seed; worth checking whether deepeval's AnthropicModel passes it through, since a non-zero temperature would explain much of this.
  • Report a per-question pass rate over repeats rather than a single pass/fail, making the instability visible instead of averaging it into a headline.

The first check should be the cheapest: whether the judge is being called with non-zero temperature. If so, most of this may be configuration rather than an inherent limit.

Found while diagnosing abstention 0/8 (map #297, PR #311).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions