You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
context-graph-eval: judge scores were attributed to the wrong questions #387
runner._judge_group paired deepeval's result.test_results with the input goldens by position (zip(goldens, result.test_results)). deepeval returns test results in completion order under async concurrency (AsyncConfig(max_concurrent=4)), so every question was scored with some other question's metrics.
Evidence
Reproduction: a fake metric with random latency, 12 cases, concurrency 4. They came back in the order q-1, q-2, q-3, q-5, q-0, q-8, …. The new regression test fails on the old code and passes on the fix.
In saved runs: 11–24 questions per 100-question run had a Contextual Recall reason quoting a different question's expected answer. This held for both a Sonnet and an OpenAI judge, and for text-search, graph-agent and hybrid runs alike.
Impact
Run totals were roughly right: a permutation within each judged group (abstention and non-abstention are judged separately) preserves the count. They aren't exact, because enforce_retrieval_floor then zeroed scores against the wrong questions' payloads.
Part of #297
What was wrong
runner._judge_grouppaired deepeval'sresult.test_resultswith the input goldens by position (zip(goldens, result.test_results)). deepeval returns test results in completion order under async concurrency (AsyncConfig(max_concurrent=4)), so every question was scored with some other question's metrics.Evidence
q-1, q-2, q-3, q-5, q-0, q-8, …. The new regression test fails on the old code and passes on the fix.reasonquoting a different question's expected answer. This held for both a Sonnet and an OpenAI judge, and for text-search, graph-agent and hybrid runs alike.Impact
enforce_retrieval_floorthen zeroed scores against the wrong questions' payloads.Fix
PR #373,
7aa430b: test cases carry the golden'sname, and results are matched by it.