Part of #390
Question
Every measurement so far is on LongMemEval: personal-assistant chat, one user, facts about the user. The context graph's own target is coding-agent sessions. Is there a corpus, even a small hand-labelled one from real Claude Code sessions, on which the typed graph's lanes can be measured? Does the hand vocabulary even fit it?
Output
A corpus proposal with a handful of questions and gold answers, or a finding that the graph lanes can't be judged on coding sessions yet.
Part of #390
Question
Every measurement so far is on LongMemEval: personal-assistant chat, one user, facts about the user. The context graph's own target is coding-agent sessions. Is there a corpus, even a small hand-labelled one from real Claude Code sessions, on which the typed graph's lanes can be measured? Does the hand vocabulary even fit it?
Output
A corpus proposal with a handful of questions and gold answers, or a finding that the graph lanes can't be judged on coding sessions yet.