Skip to content

context-graph-eval: answer with the question date, compute over listed evidence, turns first - #405

Merged
antejavor merged 2 commits into
feat/eval-hybrid-retrievalfrom
feat/398-answer-prompt
Oct 2, 2026
Merged

antejavor merged 2 commits into
feat/eval-hybrid-retrievalfrom
feat/398-answer-prompt

Conversation

@antejavor

@antejavor antejavor commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Part of #390. Closes #398. Stacked on #389 (base feat/eval-hybrid-retrieval); GitHub retargets it to main when #389 merges.

Why

On the rebuilt graph (#392 + #396) and the answer gate (#402), the full hybrid scored 60/100. Most misses with the evidence already in context were answering errors:

  • No question date. The eval never passed it unless --question-date was set, and it wasn't, so "how many weeks ago" answers counted from an invented "today" (2024-06-13).
  • Miscounting. 4 weddings for 3, 4 plants for 3, 152 game hours for 140; "two weeks" and "a week and a half" never summed.
  • Recommendations refused. All 6 preference questions answered "not in memory" under the "answer only from the rows" rule.
  • Turns read past. About 45 fact rows came before the turns, which are the answer store.

What changes

  • The question date always reaches the answer prompt. --question-date and RunPlan.question_date are removed. LongMemEval supplies it with every question, and in a real session it is simply now.
  • The shared answer prompt (retrieval.answer_prompt, used by every strategy):
    • relative dates inside a row ("today", "yesterday") are relative to that row's date;
    • for counts, totals and intervals, list each item with its date, count an item mentioned twice once, then compute;
    • when rows disagree, the most recent one is current;
    • recommendation questions get suggestions built from the user's stated preferences;
    • a question about something the rows never mention, or with a premise they don't show, gets "not in memory", not an answer to a related question.
  • Hybrid puts turns before facts, both in time order.

Measured

One full-hybrid run per variant on the rebuilt graph (port 7735, 100 questions, 5 sessions each), Sonnet 4.5 judge, answer gate:

type before r1 r2 (this PR)
temporal 10/26 20 22
multi-session 12/26 16 15
knowledge-update 13/16 14 15
single-session-user 14/15 15 14
single-session-assistant 11/11 10 10
preference 0/6 0 0
abstention 7/8 5 7
total 60 75 76

r1 lacked the false-premise rule, and abstention fell to 5/8: told to list and compute, the model answered a nearby question. r2 adds that rule and holds 7/8. Between r1 and r2, 5 questions flipped each way at the threshold, so ±4 per type is run-to-run noise.

Preference stays 0/6, now because of the judge, not the answers. The answers recommend what the question asks for (Premiere Pro resources for a user asking about its advanced settings; MICCAI and ISBI for a medical-imaging researcher), but the Coverage rubric checks for facts and the expected output is a description of a good answer. A preference rubric is filed as #404, kept out of this PR so the prompt and the judge don't change together.

Tests

context-graph/eval: 263 passed, 5 skipped against a real Memgraph; ruff and ty clean. New: the runner always passes the question date to the answer prompt; hybrid rows are turns then facts, each in time order.

Update: recommendation questions are never refused (fbb5e0f)

Re-judging with #406's preference rubric showed r2's false-premise rule also caught recommendation questions: "recommend a conference" names no conference in memory, so two answers fell back to "not in memory". Recommendation questions now always recommend, from the preferences the rows show plus general knowledge; "not in memory" stays for factual questions.

One full-hybrid run of this PR with #406's rubric (same graph and judge, answer gate): 80/100. Temporal 22/26, knowledge-update 15/16, single-session-user 14/15, multi-session 13/26, single-session-assistant 10/11, preference 6/6, abstention 8/8.

…d evidence, turns first

On the rebuilt graph, most misses with the evidence already retrieved were
answering errors: no question date, so "how many days ago" counted from an
invented today; counts and totals miscounted or never summed; recommendation
questions refused by an "only from the rows" rule; and turns placed after
~45 fact rows.

- The question date always reaches the answer prompt; --question-date and
  RunPlan.question_date are removed.
- The shared answer prompt says relative dates in a row are relative to that
  row, to list items with dates and dedupe before computing, that the latest
  row wins a conflict, to recommend from the user's stated preferences, and to
  decline a question about something the rows never mention instead of
  answering a related one.
- Hybrid puts turns before facts, both in time order.

One full-hybrid run on the rebuilt graph, answer gate: 60 -> 76/100
(temporal 10 -> 22, multi-session 12 -> 15, abstention 7/8 held).

Closes #398
…f a specific item

The false-premise rule also caught recommendation questions: 'recommend a
conference' names no conference in memory, so the answer fell back to 'not in
memory' and threw away the user's stated interests. Recommendation questions
now always recommend, from the rows' preferences plus general knowledge.
@antejavor
antejavor merged commit 066b2d4 into feat/eval-hybrid-retrieval Oct 2, 2026
16 checks passed
antejavor added a commit that referenced this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant