Skip to content

context-graph-eval: judge scores were attributed to the wrong questions #387

Description

@antejavor

Part of #297

What was wrong

runner._judge_group paired deepeval's result.test_results with the input goldens by position (zip(goldens, result.test_results)). deepeval returns test results in completion order under async concurrency (AsyncConfig(max_concurrent=4)), so every question was scored with some other question's metrics.

Evidence

  • Reproduction: a fake metric with random latency, 12 cases, concurrency 4. They came back in the order q-1, q-2, q-3, q-5, q-0, q-8, …. The new regression test fails on the old code and passes on the fix.
  • In saved runs: 11–24 questions per 100-question run had a Contextual Recall reason quoting a different question's expected answer. This held for both a Sonnet and an OpenAI judge, and for text-search, graph-agent and hybrid runs alike.

Impact

  • Run totals were roughly right: a permutation within each judged group (abstention and non-abstention are judged separately) preserves the count. They aren't exact, because enforce_retrieval_floor then zeroed scores against the wrong questions' payloads.
  • Every per-question and per-type result was wrong. That includes eval: the judge scores identical answers differently, which is what makes the passing set unstable #324's per-question pass rates and its "zero questions passing in all three runs" reading, which should be re-measured.

Fix

PR #373, 7aa430b: test cases carry the golden's name, and results are matched by it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions