Part of #390. Follow-up from #417.
Problem
Recall computes similarity exactly, over the asking user's own vectors (vector_search.cosine_similarity in Cypher), because vector_search.search has no filter. The cost grows linearly with one user's history. Measured on a synthetic user:
| Messages |
Recall, median |
| 183 (largest eval user) |
0.6 s |
| 10k |
0.7 s |
| 100k |
6.7 s |
At 100k messages, the turns lane takes 3.0 s, the text lane 1.8 s (1.1 s of that is fetching the user's own turn ids), and each graph lane about 1 s.
What to build
The fallback agreed in #393: global ANN vector indexes (node and edge) with an over-fetch, filtered to the user afterwards.
- Measure the over-fetch needed to keep the exact top-k on a multi-user graph, and fall back to the exact scan when the filtered result comes up short.
- Use the ANN path only above a history size where the exact scan gets slow.
- The text lane's own-turn list is the other half of the 100k cost: filter by ownership in the traversal instead of passing a list.
Done when
- Recall at 100k messages is under 1 s.
- The parity run still holds (official judge within noise of
g417-recall-r1, evidence recall not lower).
Part of #390. Follow-up from #417.
Problem
Recall computes similarity exactly, over the asking user's own vectors (
vector_search.cosine_similarityin Cypher), becausevector_search.searchhas no filter. The cost grows linearly with one user's history. Measured on a synthetic user:At 100k messages, the turns lane takes 3.0 s, the text lane 1.8 s (1.1 s of that is fetching the user's own turn ids), and each graph lane about 1 s.
What to build
The fallback agreed in #393: global ANN vector indexes (node and edge) with an over-fetch, filtered to the user afterwards.
Done when
g417-recall-r1, evidence recall not lower).