Drafted by Mycroft, the lab's synthetic AI co-founder; reviewed by @tonydzi.
Goal. Add eval/run_ablation.py: a sweep over the graph-expansion settings that scores every combination with the existing metrics and writes one results table, per question class. The point is to answer "does the graph help, for which kinds of query, at which settings" with numbers instead of one average.
Why it matters. The retrieval pipeline has five knobs that interact (below), and the repo has only ever scored two settings: vector-only and the shipped graph setting. Anyone tuning a memory layer over their own notes needs to know which knobs matter for which query class (title lookup, body lookup, "bridge" questions whose answers are linked notes, temporal questions). Other people's corpora will disagree with ours, so the runner has to work on any markdown folder.
Where.
eval/run_ablation.py: NEW. Reuse the pipeline functions exactly as eval/run_eval.py (score_gold()) does, so the sweep measures the code that ships: ba.query_sims(), ba.retrieve_candidates(sims, meta, topk=...), ba.expand_1hop(sims, meta, base, by_base, ghops=..., gmax=...), ba.rerank_candidates(ce, query, cand, meta, topn=...), with ba = sqlite_graph_memory.brain_ask. Load the encoder and reranker once; compute query_sims once per question; only the candidate-set and rerank steps vary per cell.
- The five axes, with defaults in
src/sqlite_graph_memory/brain_ask.py (TOPK_RETRIEVE = 60, GHOPS, GMAX = 15, 40, looks_like_entity()): (1) hops: 0 (vector only) and 1 (what expand_1hop does; deeper hops do not exist in the code yet, so treat 2+ as a follow-up issue, not part of this one); (2) seed cap ghops (how many top hits are expanded), e.g. 5, 15, 30; (3) neighbour cap gmax, e.g. 10, 40, 80; (4) rerank pool topk (candidates entering the reranker), e.g. 30, 60, 100; (5) gating: entity-like queries (looks_like_entity(query)) get the graph switched off, on or off for the sweep. Keep the default grid small enough to finish in a reasonable time and let --grid override it.
- Write one results table (CSV and a Markdown view) with columns: cell settings, class, n, Recall@12, MRR, nDCG@12, plus a column for the difference against the vector-only cell. Reuse
recall_at, mrr_at, ndcg_at from eval/run_eval.py.
README.md: add a short "Ablation" section under the eval chapter with the command and, once you have run it, the table from the mini-vault or your own notes folder.
tests/test_eval.py (or a new file): test the grid enumeration and the per-cell aggregation with no models.
How to check.
pip install -e '.[embeddings]'
python eval/build_gold.py <folder-of-markdown> --n-per-class 60 # after: sgm-index <folder>
python eval/run_ablation.py --list-grid # prints the cells, loads no model
python eval/run_ablation.py --gold eval/gold-<date>.jsonl --out ablation.csv
python -m pytest -q
Expected: --list-grid prints every combination and its count; the full run writes ablation.csv with one row per (cell, class). Baseline today: eval/run_ablation.py does not exist; python eval/run_eval.py --selftest prints SELFTEST PASS and python -m pytest -q gives 88 passed. If the synthetic mini-vault from issue #16 has landed, use it as the corpus so results are reproducible; otherwise any folder of markdown notes with [[wikilinks]] works.
Done when
Size. ~1 day
Ask here. https://github.com/tonydzi/sqlite-graph-memory/discussions, or comment on this issue.
Claim it by commenting "claiming this" — no permission needed, and it is yours for 7 days.
You keep the copyright to your code. No CLA, no assignment, ever. We answer every issue and PR within 48 hours, including "no, and here is why" — our silence is our bug, so ping the thread.
Full deal: CONTRIBUTING.md
Drafted by Mycroft, the lab's synthetic AI co-founder; reviewed by @tonydzi.
Goal. Add
eval/run_ablation.py: a sweep over the graph-expansion settings that scores every combination with the existing metrics and writes one results table, per question class. The point is to answer "does the graph help, for which kinds of query, at which settings" with numbers instead of one average.Why it matters. The retrieval pipeline has five knobs that interact (below), and the repo has only ever scored two settings: vector-only and the shipped graph setting. Anyone tuning a memory layer over their own notes needs to know which knobs matter for which query class (title lookup, body lookup, "bridge" questions whose answers are linked notes, temporal questions). Other people's corpora will disagree with ours, so the runner has to work on any markdown folder.
Where.
eval/run_ablation.py: NEW. Reuse the pipeline functions exactly aseval/run_eval.py(score_gold()) does, so the sweep measures the code that ships:ba.query_sims(),ba.retrieve_candidates(sims, meta, topk=...),ba.expand_1hop(sims, meta, base, by_base, ghops=..., gmax=...),ba.rerank_candidates(ce, query, cand, meta, topn=...), withba=sqlite_graph_memory.brain_ask. Load the encoder and reranker once; computequery_simsonce per question; only the candidate-set and rerank steps vary per cell.src/sqlite_graph_memory/brain_ask.py(TOPK_RETRIEVE = 60,GHOPS, GMAX = 15, 40,looks_like_entity()): (1) hops: 0 (vector only) and 1 (whatexpand_1hopdoes; deeper hops do not exist in the code yet, so treat 2+ as a follow-up issue, not part of this one); (2) seed capghops(how many top hits are expanded), e.g. 5, 15, 30; (3) neighbour capgmax, e.g. 10, 40, 80; (4) rerank pooltopk(candidates entering the reranker), e.g. 30, 60, 100; (5) gating: entity-like queries (looks_like_entity(query)) get the graph switched off, on or off for the sweep. Keep the default grid small enough to finish in a reasonable time and let--gridoverride it.recall_at,mrr_at,ndcg_atfromeval/run_eval.py.README.md: add a short "Ablation" section under the eval chapter with the command and, once you have run it, the table from the mini-vault or your own notes folder.tests/test_eval.py(or a new file): test the grid enumeration and the per-cell aggregation with no models.How to check.
Expected:
--list-gridprints every combination and its count; the full run writesablation.csvwith one row per (cell, class). Baseline today:eval/run_ablation.pydoes not exist;python eval/run_eval.py --selftestprintsSELFTEST PASSandpython -m pytest -qgives 88 passed. If the synthetic mini-vault from issue #16 has landed, use it as the corpus so results are reproducible; otherwise any folder of markdown notes with[[wikilinks]]works.Done when
eval/run_ablation.pysweeps all five axes (hops limited to 0 and 1) and writes one table with a row per cell and classnshown and a warning for classes under 20 questionsbrain_ask.py(no reimplemented pipeline), and the no-model parts have testsSize. ~1 day
Ask here. https://github.com/tonydzi/sqlite-graph-memory/discussions, or comment on this issue.
Claim it by commenting "claiming this" — no permission needed, and it is yours for 7 days.
You keep the copyright to your code. No CLA, no assignment, ever. We answer every issue and PR within 48 hours, including "no, and here is why" — our silence is our bug, so ping the thread.
Full deal: CONTRIBUTING.md