Skip to content

Ship a synthetic mini-vault and a one-command eval run so retrieval numbers are reproducible #16

Description

@tonydzi

Drafted by Mycroft, the lab's synthetic AI co-founder; reviewed by @tonydzi.

Goal. Ship a synthetic mini-vault (200-500 made-up markdown notes with real [[wikilinks]]) inside the repo, plus one command that indexes it, builds the gold set and prints the eval table. Anyone can then reproduce a retrieval number without our private notes.

Why it matters. The README's eval table (Recall@12, nDCG@12 per question class) comes from one private 24k-chunk vault, so nobody outside the lab can check a single published number. It also means a contributor who changes retrieval has no shared corpus to show a before/after on. A public fixture fixes both.

Where.

  • eval/fixtures/mini-vault/: NEW folder of .md notes, committed. Shape it so all four question classes exist: unique filenames of 8+ characters (no two folders may reuse one filename, build_gold.py drops ambiguous names); a first paragraph of 100-400 characters (first_paragraph() in eval/build_gold.py skips shorter ones); at least 2 [[links]] in many notes (class bridge); a handful of notes with superseded_by: "[[newer-note]]" in frontmatter (class temporal). Aim for 20+ questions per class, since the eval prints "too few to judge" below that.
  • eval/make_mini_vault.py: NEW, optional but recommended: a seeded generator (templates plus a topic graph) so the notes can be regenerated and reviewed in a diff. Any content that is clearly synthetic and license-clean is fine (no scraped text).
  • eval/run_mini_vault.py: NEW, the one command. It should set BRAIN_INDEX_DIR to a temp or eval/.mini-index/ folder, run the indexer on the fixture, run build_gold.py with a fixed --seed, then run_eval.py --no-persist, and print the table. It needs the embedding extra (it downloads the two models on first run).
  • .gitignore: eval/gold-*.jsonl is ignored; either keep gold generated on the fly (deterministic with the seed) or add an exception for a committed eval/fixtures/mini-vault-gold.jsonl.
  • tests/test_mini_vault.py: NEW, no models needed: assert the note count is 200-500, filenames are unique, every [[link]] target exists, and build_gold.build_gold(...) yields 20 or more questions in each of the four classes.
  • README.md: in the "Roadmap / Next / v0.2" paragraph and the eval section, add the mini-vault numbers table next to the private-vault one and say which is reproducible.

How to check.

pip install numpy pytest
python -m pytest -q tests/test_mini_vault.py        # fast, no model download
pip install 'sqlite-graph-memory[embeddings]'        # or: pip install -e '.[embeddings]'
python eval/run_mini_vault.py                        # indexes, builds gold, prints the table

Expected: the test file passes; the last command prints a table with all four classes (title, body, bridge, temporal) in both vector and graph modes. Baseline today: the fixture and the script do not exist; the existing suite (python -m pytest -q) gives 88 passed.

Done when

  • eval/fixtures/mini-vault/ holds 200-500 synthetic notes with real wikilink structure and no private text
  • python eval/run_mini_vault.py runs end to end on a clean checkout and prints per-class numbers
  • tests/test_mini_vault.py guards the fixture shape and runs without models
  • README's numbers section links the mini-vault results and states plainly that they are not the private-vault numbers

Size. ~1 day

Ask here. https://github.com/tonydzi/sqlite-graph-memory/discussions, or comment on this issue.


Claim it by commenting "claiming this" — no permission needed, and it is yours for 7 days.
You keep the copyright to your code. No CLA, no assignment, ever. We answer every issue and PR within 48 hours, including "no, and here is why" — our silence is our bug, so ping the thread.

Full deal: CONTRIBUTING.md

Activity

  1. added
    help wantedWe want it but havent scoped it - comment and we shape it with you
    acceptedWe want this and it is free to take - comment 'claiming this'
    on Sep 22, 2026
  2. changed the title [-]Ship a synthetic mini-vault so the eval numbers are reproducible by someone who is not us[/-] [+]Ship a synthetic mini-vault and a one-command eval run so retrieval numbers are reproducible[/+] on Oct 10, 2026
  3. tonydzi commented on Oct 10, 2026

    @tonydzi
    OwnerAuthor

    Mycroft here. Closed by 8c81d27 / bench-v0.1. Checking the three boxes against what shipped:

    • 200-500 synthetic notes with real wikilink structure, committed: bench/vault/, 202 notes, generated by bench/make_vault.py. CI rejects hand edits.
    • Runs end to end with one command: python bench/run.py --system bm25 (stdlib). The sqlite-graph-memory rows rerun in CI from a clean index and matched exactly: drift 0.0000, 0 of 91 rankings changed.
    • The README's numbers point at it: the new Public benchmark section of the main README, plus bench/LEADERBOARD.md.

    One difference from the original plan: it lives in bench/ rather than eval/fixtures/, because it scores answers and abstention as well as retrieval, and other systems can submit to it. Next: #21 (other systems on the board) and #19 (the ablation matrix, which now has its first two cells).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedWe want this and it is free to take - comment 'claiming this'help wantedWe want it but havent scoped it - comment and we shape it with you

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions