Part of #390
Problem
A question counts as covered when the weaker of its two judge scores reaches the threshold. The two scores are Contextual Recall (does the retrieved context support the expected answer) and the Coverage rubric (is the answer right). Contextual Recall can't support a computed answer such as "7 days", "17" or "2", because no single row states it. So a correct computed answer scores 0.
In the #391 run, 7 of the 55 misses had a correct answer:
gpt4_59149c77: 7 days;
gpt4_1916e0ea: 54 days;
6cb6f249: 17;
b5ef892d: 8;
80ec1f4f: 2;
d7c942c3: Yes;
aae3761f: 15 hours.
Temporal and multi-session questions are hit hardest.
Options
- Gate on the Coverage rubric only, and keep Contextual Recall as a reported retrieval signal that doesn't gate. This pairs a hard automatic gate with a reporting-only signal.
- Keep the gate, but give Contextual Recall the evidence turns instead of the expected answer.
Acceptance
Correct computed answers count as covered. Retrieval quality stays visible per question.
Part of #390
Problem
A question counts as covered when the weaker of its two judge scores reaches the threshold. The two scores are Contextual Recall (does the retrieved context support the expected answer) and the Coverage rubric (is the answer right). Contextual Recall can't support a computed answer such as "7 days", "17" or "2", because no single row states it. So a correct computed answer scores 0.
In the #391 run, 7 of the 55 misses had a correct answer:
gpt4_59149c77: 7 days;gpt4_1916e0ea: 54 days;6cb6f249: 17;b5ef892d: 8;80ec1f4f: 2;d7c942c3: Yes;aae3761f: 15 hours.Temporal and multi-session questions are hit hardest.
Options
Acceptance
Correct computed answers count as covered. Retrieval quality stays visible per question.