Repository navigation
context-graph-eval: answer with the question date, compute over listed evidence, turns first - #405
Merged
Conversation
…d evidence, turns first On the rebuilt graph, most misses with the evidence already retrieved were answering errors: no question date, so "how many days ago" counted from an invented today; counts and totals miscounted or never summed; recommendation questions refused by an "only from the rows" rule; and turns placed after ~45 fact rows. - The question date always reaches the answer prompt; --question-date and RunPlan.question_date are removed. - The shared answer prompt says relative dates in a row are relative to that row, to list items with dates and dedupe before computing, that the latest row wins a conflict, to recommend from the user's stated preferences, and to decline a question about something the rows never mention instead of answering a related one. - Hybrid puts turns before facts, both in time order. One full-hybrid run on the rebuilt graph, answer gate: 60 -> 76/100 (temporal 10 -> 22, multi-session 12 -> 15, abstention 7/8 held). Closes #398
…f a specific item The false-premise rule also caught recommendation questions: 'recommend a conference' names no conference in memory, so the answer fell back to 'not in memory' and threw away the user's stated interests. Recommendation questions now always recommend, from the rows' preferences plus general knowledge.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #390. Closes #398. Stacked on #389 (base
feat/eval-hybrid-retrieval); GitHub retargets it tomainwhen #389 merges.Why
On the rebuilt graph (#392 + #396) and the answer gate (#402), the full hybrid scored 60/100. Most misses with the evidence already in context were answering errors:
--question-datewas set, and it wasn't, so "how many weeks ago" answers counted from an invented "today" (2024-06-13).What changes
--question-dateandRunPlan.question_dateare removed. LongMemEval supplies it with every question, and in a real session it is simply now.retrieval.answer_prompt, used by every strategy):Measured
One full-hybrid run per variant on the rebuilt graph (port 7735, 100 questions, 5 sessions each), Sonnet 4.5 judge, answer gate:
r1 lacked the false-premise rule, and abstention fell to 5/8: told to list and compute, the model answered a nearby question. r2 adds that rule and holds 7/8. Between r1 and r2, 5 questions flipped each way at the threshold, so ±4 per type is run-to-run noise.
Preference stays 0/6, now because of the judge, not the answers. The answers recommend what the question asks for (Premiere Pro resources for a user asking about its advanced settings; MICCAI and ISBI for a medical-imaging researcher), but the Coverage rubric checks for facts and the expected output is a description of a good answer. A preference rubric is filed as #404, kept out of this PR so the prompt and the judge don't change together.
Tests
context-graph/eval: 263 passed, 5 skipped against a real Memgraph; ruff and ty clean. New: the runner always passes the question date to the answer prompt; hybrid rows are turns then facts, each in time order.Update: recommendation questions are never refused (
fbb5e0f)Re-judging with #406's preference rubric showed r2's false-premise rule also caught recommendation questions: "recommend a conference" names no conference in memory, so two answers fell back to "not in memory". Recommendation questions now always recommend, from the preferences the rows show plus general knowledge; "not in memory" stays for factual questions.
One full-hybrid run of this PR with #406's rubric (same graph and judge, answer gate): 80/100. Temporal 22/26, knowledge-update 15/16, single-session-user 14/15, multi-session 13/26, single-session-assistant 10/11, preference 6/6, abstention 8/8.