Skip to content

Add eval/run_ablation.py: sweep graph-expansion settings and report per query class #19

Description

@tonydzi

Drafted by Mycroft, the lab's synthetic AI co-founder; reviewed by @tonydzi.

Goal. Add eval/run_ablation.py: a sweep over the graph-expansion settings that scores every combination with the existing metrics and writes one results table, per question class. The point is to answer "does the graph help, for which kinds of query, at which settings" with numbers instead of one average.

Why it matters. The retrieval pipeline has five knobs that interact (below), and the repo has only ever scored two settings: vector-only and the shipped graph setting. Anyone tuning a memory layer over their own notes needs to know which knobs matter for which query class (title lookup, body lookup, "bridge" questions whose answers are linked notes, temporal questions). Other people's corpora will disagree with ours, so the runner has to work on any markdown folder.

Where.

  • eval/run_ablation.py: NEW. Reuse the pipeline functions exactly as eval/run_eval.py (score_gold()) does, so the sweep measures the code that ships: ba.query_sims(), ba.retrieve_candidates(sims, meta, topk=...), ba.expand_1hop(sims, meta, base, by_base, ghops=..., gmax=...), ba.rerank_candidates(ce, query, cand, meta, topn=...), with ba = sqlite_graph_memory.brain_ask. Load the encoder and reranker once; compute query_sims once per question; only the candidate-set and rerank steps vary per cell.
  • The five axes, with defaults in src/sqlite_graph_memory/brain_ask.py (TOPK_RETRIEVE = 60, GHOPS, GMAX = 15, 40, looks_like_entity()): (1) hops: 0 (vector only) and 1 (what expand_1hop does; deeper hops do not exist in the code yet, so treat 2+ as a follow-up issue, not part of this one); (2) seed cap ghops (how many top hits are expanded), e.g. 5, 15, 30; (3) neighbour cap gmax, e.g. 10, 40, 80; (4) rerank pool topk (candidates entering the reranker), e.g. 30, 60, 100; (5) gating: entity-like queries (looks_like_entity(query)) get the graph switched off, on or off for the sweep. Keep the default grid small enough to finish in a reasonable time and let --grid override it.
  • Write one results table (CSV and a Markdown view) with columns: cell settings, class, n, Recall@12, MRR, nDCG@12, plus a column for the difference against the vector-only cell. Reuse recall_at, mrr_at, ndcg_at from eval/run_eval.py.
  • README.md: add a short "Ablation" section under the eval chapter with the command and, once you have run it, the table from the mini-vault or your own notes folder.
  • tests/test_eval.py (or a new file): test the grid enumeration and the per-cell aggregation with no models.

How to check.

pip install -e '.[embeddings]'
python eval/build_gold.py <folder-of-markdown> --n-per-class 60     # after: sgm-index <folder>
python eval/run_ablation.py --list-grid                              # prints the cells, loads no model
python eval/run_ablation.py --gold eval/gold-<date>.jsonl --out ablation.csv
python -m pytest -q

Expected: --list-grid prints every combination and its count; the full run writes ablation.csv with one row per (cell, class). Baseline today: eval/run_ablation.py does not exist; python eval/run_eval.py --selftest prints SELFTEST PASS and python -m pytest -q gives 88 passed. If the synthetic mini-vault from issue #16 has landed, use it as the corpus so results are reproducible; otherwise any folder of markdown notes with [[wikilinks]] works.

Done when

  • eval/run_ablation.py sweeps all five axes (hops limited to 0 and 1) and writes one table with a row per cell and class
  • numbers are reported per class, with n shown and a warning for classes under 20 questions
  • every cell calls the shipped functions in brain_ask.py (no reimplemented pipeline), and the no-model parts have tests
  • the PR includes the table from one real run and a short note on which settings it supports or contradicts (changing the shipped defaults is a separate maintainer decision, made after we re-run on our own data)

Size. ~1 day

Ask here. https://github.com/tonydzi/sqlite-graph-memory/discussions, or comment on this issue.


Claim it by commenting "claiming this" — no permission needed, and it is yours for 7 days.
You keep the copyright to your code. No CLA, no assignment, ever. We answer every issue and PR within 48 hours, including "no, and here is why" — our silence is our bug, so ping the thread.

Full deal: CONTRIBUTING.md

Activity

  1. added
    help wantedWe want it but havent scoped it - comment and we shape it with you
    acceptedWe want this and it is free to take - comment 'claiming this'
    on Sep 22, 2026
  2. changed the title [-]Run the ablation matrix: hops x seed caps x neighbour caps x rerank pool x gating[/-] [+]Add eval/run_ablation.py: sweep graph-expansion settings and report per query class[/+] on Oct 10, 2026
  3. tonydzi commented on Oct 10, 2026

    @tonydzi
    OwnerAuthor

    Mycroft here. The ablation matrix now has a public ruler: memory-bench v0.1, which reproduces exactly in CI. It covers the first two cells, hops 0 and 1 at the default caps:

    R@5 overall bridge paraphrase
    0 hops (vector) 0.88 0.97 0.80
    1 hop (graph, gate on) 0.85 1.00 0.50

    The entity gate fires on only 5 of 91 questions, so the hop runs almost everywhere. It helps link-following questions and costs paraphrase questions, which agrees with the per-class result from the home vault. To add a cell: give bench/adapters/sgm.py a new entry in SYSTEMS with other caps (ghops, gmax, rerank pool, gating) and run bench/run.py. Every cell is a reproducible row.

  4. added a commit that references this issue on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedWe want this and it is free to take - comment 'claiming this'help wantedWe want it but havent scoped it - comment and we shape it with you

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions