Part of #390. From the #399 grilling. Depends on #413.
What
Rebuild the eval graph once with the new extraction, then measure:
- graph quality (
graph-quality) against the baseline taken on today's graph;
- one full-hybrid run on LongMemEval's judge as the no-regression check (today 84–85/100; single-session-assistant 10/11 is the one to watch, since assistant-turn edges shrink).
Acceptance
Both reports, compared with the baselines.
Part of #390. From the #399 grilling. Depends on #413.
What
Rebuild the eval graph once with the new extraction, then measure:
graph-quality) against the baseline taken on today's graph;Acceptance
Both reports, compared with the baselines.