You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Map: Hybrid retrieval — the typed graph as an index into turns #390
A retrieval design for the context graph that uses the typed graph as an index into turn text. It's implemented as the sessions-graph read path. Target: 80–90% correct end to end on a trustworthy judge, with an understanding of what each stage contributes and why it works.
The eval side exists as a starting point in #389: --retrieval-strategy hybrid, with lanes for turn vectors, turn full text, entities and their edges, facts by source sentence, and the user's facts across sessions. This map decides what of it becomes the product, how, and on what evidence.
Why this map exists
Map #344 made the graph hold typed, time-anchored facts (merged in #373). Reading it back has so far shown:
Typed edges are a poor answer store. On 100 LongMemEval questions with 5 sessions each, answering from every typed edge of the right sessions scored 21/100; answering from their full text scored 48/100. An LLM agent writing Cypher over the typed schema scored 12–15, below plain text search (33–34).
Typed edges are a usable index. Hybrid retrieval, with the graph finding facts and their source turns, scored 40–44 against 37–38 for turns alone, across three runs each. It helped most on knowledge-update questions, which reached 14/16 on a cleaner graph.
Domain:context-graph/eval for measurement; unstructured2graph and sessions-graph for what ships.
Skills:/grilling, /domain-modeling and /research, as on the other maps.
Decision criterion: end-to-end coverage of full hybrid against the current baseline (45/100, Does the graph lift hybrid retrieval on corrected scoring? #391), same graph and judge; target 80–90%. Driven by failure analysis: which questions are missed and why (retrieval, graph model, or answering). Ablations explain contributions; they don't gate. Text search and the Cypher graph agent are not re-run, since they're not the frontier.
Baseline and criterion (Does the graph lift hybrid retrieval on corrected scoring? #391): full hybrid on the round-2 graph scores 45/100 on corrected scoring (knowledge-update 14/16, single-session-user 11/15, single-session-assistant 9/11, temporal 7/26, multi-session 4/26, preference 0/6; abstention 7/8). The turns-only ablation no longer gates; the next step is failure analysis of the 55 misses.
Embeddings (Where embeddings live, and what gets embedded #393): vectors live in Memgraph, computed by MAGE embeddings.text; bge-small-en-v1.5 by default; messages, entity names and edge text; messages embedded at session end; exact cosine over the user's own rows.
Read path (The retrieval read path: API and what the answerer receives #394): recall(question) returns context, the harness model answers; served over MCP by agent-context-graph mcp (tools registered through the agent_context_graph.tools entry point), plus agent-context-graph recall [--json]; own memory only, user from the config file; widths from [recall]; a SessionStart line points the model at it. MCP results are text only: Claude Code and Codex both show the model structuredContent instead of the text when both are sent.
Not pursued, since the target is met; reopen only if a new measurement calls for it:
Whether retrieval should iterate (answer, find what's missing, retrieve again) or stay single-pass.
How the lanes weight and budget against each other.
Answering over aggregates: largely closed by the answer prompt (multi-session 4/26 → 13–15/26, temporal 7/26 → 22/26); the remaining misses split between partial retrieval and answer noise.
Destination
A retrieval design for the context graph that uses the typed graph as an index into turn text. It's implemented as the sessions-graph read path. Target: 80–90% correct end to end on a trustworthy judge, with an understanding of what each stage contributes and why it works.
The eval side exists as a starting point in #389:
--retrieval-strategy hybrid, with lanes for turn vectors, turn full text, entities and their edges, facts by source sentence, and the user's facts across sessions. This map decides what of it becomes the product, how, and on what evidence.Why this map exists
Map #344 made the graph hold typed, time-anchored facts (merged in #373). Reading it back has so far shown:
Notes
context-graph/evalfor measurement;unstructured2graphandsessions-graphfor what ships./grilling,/domain-modelingand/research, as on the other maps.--skip-reconcile; re-extract only when extraction changes.Decisions so far
Segment.source_id; asourceslist onMENTIONED_IN; one edge per turn withsource_id,textandrole;link_turnsdeleted; the eval graph rebuilt.embeddings.text; bge-small-en-v1.5 by default; messages, entity names and edge text; messages embedded at session end; exact cosine over the user's own rows.recall(question)returns context, the harness model answers; served over MCP byagent-context-graph mcp(tools registered through theagent_context_graph.toolsentry point), plusagent-context-graph recall [--json]; own memory only, user from the config file; widths from[recall]; a SessionStart line points the model at it. MCP results are text only: Claude Code and Codex both show the modelstructuredContentinstead of the text when both are sent.recall()core shared by the eval, the MCP server in both plugins, and assistant replies recorded at turn end. Parity run on the shipped code: 89/100 on the official judge, evidence recall 79/92; recall takes 0.6 s on eval users, 0.7 s at 10k messages, 6.7 s at 100k. The 80–90% target is met.Not yet specified
Not pursued, since the target is met; reopen only if a new measurement calls for it:
Out of scope