Repository navigation
context-graph: sessions-graph recall(), the read path the eval benchmarks - #424
Merged
Merged
Conversation
…arks sessions_graph/recall.py moves the eval's hybrid retrieval into the product: five lanes (turns, text, entities, facts, user_facts) as Cypher over the asking user's own rows, with similarity computed exactly in Memgraph over the stored vectors. A fact or entity mention is the user's when its source turn is; chunks are shared across users. Without MAGE, recall runs text search plus the facts of the turns it found and says so. Recalled.lines() is what the answerer reads; render(today) adds the reading rules; to_json() returns the data. - Edges are embedded as "<head> <type> <tail>. <sentence>", as measured. - The eval's hybrid strategy calls recall and keeps its own answerer; the host embedder and .npz index are gone, and the runner embeds anything missing before the first question. - sessions-graph tests point context-graph's config lookup at an empty file, so the CLI tests never read the host's config. Parity run on the rebuilt eval graph, vectors computed in Memgraph: 89/100 on LongMemEval's judge (84 before), evidence 79/92 (77); 9 of the 11 flips had identical evidence recall, so this is parity within answer noise. Recall takes 0.6 s on eval users, 0.7 s at 10k messages and 6.7 s at 100k; the ANN fallback is #423. Closes #417.
This was referenced Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #417. Part of map #390 (#415 → #416 → #417 → #418). The ANN follow-up is #423.
What
sessions_graph/recall.py:vector_search.cosine_similarityover the vectors from sessions-graph: embed messages, entities and edges in Memgraph #416).Recalled.lines()is exactly what the eval's answerer reads.render(today)puts the reading rules (fromanswer_prompt) and the date before the rows, which is what agent-context-graph: MCP server with the recall tool #418's MCP tool will return.to_json()returns the same data.RecallConfig.from_mappingcarries the[recall]overrides (agent-context-graph: MCP server with the recall tool #418 reads them from the config file).SessionsGraph.recall()is the entry point, andsetup()creates the message text index.embeddings.py: edges are now embedded as"<head> <type> <tail>. <sentence>", which is what the benchmark measured.hybrid.pyis nowrecall, then the eval's ownanswer_prompt. The host embedder and.npzindex are removed (about 570 lines).sessions-graphtests point context-graph's config lookup at an empty file. Before this, the CLI tests read the developer's real config.user_factspicks among the asking user's own relation types rather than every type in the graph.Parity gate
The rebuilt eval graph on port 7735 was embedded inside Memgraph (502 sessions in 21 min, none failed). Memgraph's bge-small vectors match the host embedder exactly (cosine 1.0).
g409-official-r1(host index)g417-recall-r1(this PR)Of the 11 questions that flipped, 9 had identical evidence recall in both runs: that's answer noise, so this is parity, not a gain.
Latency
At 100k: the turns lane takes 3.0 s, the text lane 1.8 s, and each graph lane about 1 s. Accepted for now; #423 tracks the ANN fallback.
Tests
sessions-graphsessions-graph: 9 end-to-end recall tests on real MAGE:user_factsgathers across sessions;