Part of #390. Found while measuring #398.
Problem
Preference questions score 0/6 although the answers now do what the question asks. Their expected output is a description of what a good answer looks like ("The user would prefer suggestions of Sony-compatible accessories … may not prefer other brands"), not a set of facts. The Coverage rubric checks that every fact in the expected output appears in the answer, so a good recommendation scores 0.2–0.6.
From the #398 run (g398-prompt-r1), with the prompt now allowed to recommend:
8a2466db (video editing resources): recommends Adobe Premiere Pro tutorials because the user asked about its advanced settings — 0.6.
75832dbd (publications/conferences): MICCAI and ISBI, from the user's interest in medical image analysis — 0.6.
32260d93 (what to watch): stand-up specials known for storytelling — 0.6.
06878be2 (photography accessories): uses the user's Sony A7R IV setup — 0.2.
This is the same mismatch abstention questions had before they got their own rubric (scoring.build_metrics).
What to do
Give preference questions their own GEval rubric, as abstention has: does the answer recommend things that fit the preferences the expected output describes, and avoid what it says the user would not want? LongMemEval grades this question type with its own rubric too; check its wording and stay close to it so scores stay comparable with published numbers.
Acceptance
A recommendation that uses the user's stated preferences passes; a generic or "not in memory" answer fails. Re-score from one full-hybrid run.
Part of #390. Found while measuring #398.
Problem
Preference questions score 0/6 although the answers now do what the question asks. Their expected output is a description of what a good answer looks like ("The user would prefer suggestions of Sony-compatible accessories … may not prefer other brands"), not a set of facts. The Coverage rubric checks that every fact in the expected output appears in the answer, so a good recommendation scores 0.2–0.6.
From the #398 run (
g398-prompt-r1), with the prompt now allowed to recommend:8a2466db(video editing resources): recommends Adobe Premiere Pro tutorials because the user asked about its advanced settings — 0.6.75832dbd(publications/conferences): MICCAI and ISBI, from the user's interest in medical image analysis — 0.6.32260d93(what to watch): stand-up specials known for storytelling — 0.6.06878be2(photography accessories): uses the user's Sony A7R IV setup — 0.2.This is the same mismatch abstention questions had before they got their own rubric (
scoring.build_metrics).What to do
Give preference questions their own GEval rubric, as abstention has: does the answer recommend things that fit the preferences the expected output describes, and avoid what it says the user would not want? LongMemEval grades this question type with its own rubric too; check its wording and stay close to it so scores stay comparable with published numbers.
Acceptance
A recommendation that uses the user's stated preferences passes; a generic or "not in memory" answer fails. Re-score from one full-hybrid run.