Skip to content

context-graph: sessions-graph recall(), the read path the eval benchmarks - #424

Merged
antejavor merged 1 commit into
mainfrom
feat/417-recall
Oct 5, 2026
Merged

antejavor merged 1 commit into
mainfrom
feat/417-recall

Conversation

@antejavor

Copy link
Copy Markdown
Contributor

Closes #417. Part of map #390 (#415 → #416 → #417 → #418). The ANN follow-up is #423.

What

  • sessions_graph/recall.py:
    • Lanes: the eval's five hybrid lanes (turns, text, entities, facts, user_facts) moved into the product.
      • Each runs as Cypher over the asking user's own rows, with similarity computed exactly in Memgraph (vector_search.cosine_similarity over the vectors from sessions-graph: embed messages, entities and edges in Memgraph #416).
      • A fact, or an entity mention, is the user's when its source turn is, because chunks are content-addressed and shared across users.
    • Result: Recalled.lines() is exactly what the eval's answerer reads. render(today) puts the reading rules (from answer_prompt) and the date before the rows, which is what agent-context-graph: MCP server with the recall tool #418's MCP tool will return. to_json() returns the same data.
    • Without MAGE: text search plus the facts of the turns it found, and the result says the vector lanes are off.
    • Config: RecallConfig.from_mapping carries the [recall] overrides (agent-context-graph: MCP server with the recall tool #418 reads them from the config file). SessionsGraph.recall() is the entry point, and setup() creates the message text index.
  • embeddings.py: edges are now embedded as "<head> <type> <tail>. <sentence>", which is what the benchmark measured.
  • Eval:
    • hybrid.py is now recall, then the eval's own answer_prompt. The host embedder and .npz index are removed (about 570 lines).
    • The runner embeds anything missing before the first question, and fails the run if Memgraph can't embed.
  • Test isolation: sessions-graph tests point context-graph's config lookup at an empty file. Before this, the CLI tests read the developer's real config.
  • One deliberate difference from the eval: user_facts picks among the asking user's own relation types rather than every type in the graph.

Parity gate

The rebuilt eval graph on port 7735 was embedded inside Memgraph (502 sessions in 21 min, none failed). Memgraph's bge-small vectors match the host embedder exactly (cosine 1.0).

Run LongMemEval judge Rubrics All evidence retrieved
g409-official-r1 (host index) 84/100 84/100 77/92
g417-recall-r1 (this PR) 89/100 87/100 79/92

Of the 11 questions that flipped, 9 had identical evidence recall in both runs: that's answer noise, so this is parity, not a gain.

Latency

History Median
Largest eval user (183 messages) 0.6 s
Synthetic, 10k messages 0.7 s
Synthetic, 100k messages 6.7 s

At 100k: the turns lane takes 3.0 s, the text lane 1.8 s, and each graph lane about 1 s. Accepted for now; #423 tracks the ANN fallback.

Tests

Suite Result
sessions-graph 87 passed, 1 skipped
eval 276 passed, 5 skipped
  • New in sessions-graph: 9 end-to-end recall tests on real MAGE:
    • the facts lane reaches its source turn;
    • turns-only mode;
    • user_facts gathers across sessions;
    • per-user scoping on a shared entity;
    • ordering;
    • the degraded mode;
    • rendering;
    • an empty memory;
    • config validation.
  • Eval: the hybrid tests now cover only the wrapper: the graph is embedded once, and the answerer reads exactly recall's rows.

…arks

sessions_graph/recall.py moves the eval's hybrid retrieval into the
product: five lanes (turns, text, entities, facts, user_facts) as Cypher
over the asking user's own rows, with similarity computed exactly in
Memgraph over the stored vectors. A fact or entity mention is the user's
when its source turn is; chunks are shared across users. Without MAGE,
recall runs text search plus the facts of the turns it found and says so.
Recalled.lines() is what the answerer reads; render(today) adds the
reading rules; to_json() returns the data.

- Edges are embedded as "<head> <type> <tail>. <sentence>", as measured.
- The eval's hybrid strategy calls recall and keeps its own answerer; the
  host embedder and .npz index are gone, and the runner embeds anything
  missing before the first question.
- sessions-graph tests point context-graph's config lookup at an empty
  file, so the CLI tests never read the host's config.

Parity run on the rebuilt eval graph, vectors computed in Memgraph:
89/100 on LongMemEval's judge (84 before), evidence 79/92 (77); 9 of
the 11 flips had identical evidence recall, so this is parity within
answer noise. Recall takes 0.6 s on eval users, 0.7 s at 10k messages
and 6.7 s at 100k; the ANN fallback is #423.

Closes #417.
@antejavor
antejavor merged commit bde67f9 into main Oct 5, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

sessions-graph: recall() core, shared by the eval

1 participant