Skip to content

context-graph-eval: judge preference questions on a preference rubric - #406

Merged
antejavor merged 1 commit into
mainfrom
feat/404-preference-rubric
Oct 2, 2026
Merged

antejavor merged 1 commit into
mainfrom
feat/404-preference-rubric

Conversation

@antejavor

Copy link
Copy Markdown
Contributor

Part of #390. Closes #404.

Why

Preference questions ("Can you suggest a hotel for my trip to Miami?") have an expected output that describes a good personalized answer, e.g. "The user would prefer suggestions of Sony-compatible accessories or high-quality photography gear … may not prefer other brands." The Coverage rubric checks that every fact in the expected output is present, so recommendations that did exactly what the description asks scored 0.2–0.6: 0/6. Abstention had the same mismatch before it got its own rubric.

What changes

  • A Preference GEval rubric for single-session-preference questions. It follows LongMemEval's own judge for this type ("The model does not need to reflect all the points in the rubric. The response is correct as long as it recalls and utilizes the user's personal information correctly."), so scores stay comparable with published numbers. A generic answer, one that ignores the user's preferences, or "not in memory" fails.
  • scoring.rubric_for routes each golden to answer, abstention or preference; build_metrics(judge, rubric=...) replaces the abstention flag, and the runner judges each rubric in its own pass.
  • Contextual Recall stays beside Preference as the reported (not gating) retrieval signal, as for every non-abstention question since context-graph-eval: gate coverage on the answer rubric, report contextual recall #402.

Evidence

Re-judged the six preference answers saved from three runs with the new rubric (Sonnet 4.5, minimal effort), without re-running retrieval:

answers from passed
rebuilt-graph run, all "not in memory" 0/6 (all 0.00)
#405 first run 4/6 (personalized answers 0.8–0.9)
#405 second run 3/6

The rubric isn't lenient: it fails refusals and generic answers, and passes recommendations that use what the user said. The drop from the first to the second run is a prompt conflict in #405 (its false-premise rule also caught recommendation questions); it's fixed there.

Tests

context-graph/eval: 256 passed, 5 skipped against a real Memgraph; ruff and ty clean. New: routing to the right rubric, the preference metric's makeup, one judge pass per rubric.

A preference question's expected output describes a good personalized answer
("would prefer Sony-compatible accessories"), not facts, so the Coverage
rubric's every-fact check scored recommendations that did exactly that at
0.2-0.6: 0/6.

- A Preference GEval rubric, following LongMemEval's own judge for this type:
  the answer must recall and use the user's personal information; it need not
  reflect every point; a generic answer or "not in memory" fails.
- scoring.rubric_for routes each golden to answer, abstention or preference;
  the runner judges each rubric in its own pass. Contextual Recall stays
  beside Preference as the reported retrieval signal.

Re-judging saved answers: six "not in memory" answers score 0.00; the
personalized recommendations from the #398 runs score 0.8-0.9.

Closes #404
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval: a preference rubric for recommendation questions

1 participant