Skip to content

Eval: Contextual Recall fails correct computed answers #397

Description

@antejavor

Part of #390

Problem

A question counts as covered when the weaker of its two judge scores reaches the threshold. The two scores are Contextual Recall (does the retrieved context support the expected answer) and the Coverage rubric (is the answer right). Contextual Recall can't support a computed answer such as "7 days", "17" or "2", because no single row states it. So a correct computed answer scores 0.

In the #391 run, 7 of the 55 misses had a correct answer:

  • gpt4_59149c77: 7 days;
  • gpt4_1916e0ea: 54 days;
  • 6cb6f249: 17;
  • b5ef892d: 8;
  • 80ec1f4f: 2;
  • d7c942c3: Yes;
  • aae3761f: 15 hours.

Temporal and multi-session questions are hit hardest.

Options

  • Gate on the Coverage rubric only, and keep Contextual Recall as a reported retrieval signal that doesn't gate. This pairs a hard automatic gate with a reporting-only signal.
  • Keep the gate, but give Contextual Recall the evidence turns instead of the expected answer.

Acceptance

Correct computed answers count as covered. Retrieval quality stays visible per question.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions