Part of #390
Problem
Answering errors in the #391 run, where all the evidence was already in context:
- A turn's timestamp is its "today". In
gpt4_1d4ab0c9, 0db4c65d and gpt4_1d80365e, the turns say "…today" and carry their dates, and the answer was "not in memory".
- Facts crowd out turns. About 45 FACT rows come before the TURN rows, so evidence sits deep in the context. In
89527b6b, "blue" is at character 1,259 of a retrieved turn and was missed.
- Aggregation.
e831120c found "two weeks" and "a week and a half" but didn't add them up.
- Preference questions scored 0 of 6. The prompt says to answer only from the rows, briefly, or say "not in memory". The gold answers are personalised recommendations built from what the user said.
- The question date. Without
--question-date, 8 "how long ago" questions guessed today's date. Decide whether the eval always passes it.
Acceptance
A single full-hybrid run on the rebuilt graph, against the re-measured baseline.
Part of #390
Problem
Answering errors in the #391 run, where all the evidence was already in context:
gpt4_1d4ab0c9,0db4c65dandgpt4_1d80365e, the turns say "…today" and carry their dates, and the answer was "not in memory".89527b6b, "blue" is at character 1,259 of a retrieved turn and was missed.e831120cfound "two weeks" and "a week and a half" but didn't add them up.--question-date, 8 "how long ago" questions guessed today's date. Decide whether the eval always passes it.Acceptance
A single full-hybrid run on the rebuilt graph, against the re-measured baseline.