Repository navigation
Releases: tonydzi/sqlite-graph-memory
Release list
memory-bench v0.1
memory-bench v0.1 is a public benchmark for retrieval from an agent's markdown memory. It lives in bench/.
- 202 synthetic notes about a fictional lab, linked with
[[wikilinks]], with no personal data. - 91 questions in seven classes: fact, bridge, paraphrase, paraphrase+bridge, temporal, absent entity, absent attribute. 22 of them have no answer in the vault.
- Metrics: recall@k on support notes, verbatim citation, abstaining on a missing fact.
python bench/run.py --system bm25runs end to end with nothing installed.
Baseline, with its weak spots (leaderboard):
| system | R@5 | answer | abstains when absent |
|---|---|---|---|
| sqlite-graph-memory (vector) | 0.88 | 0.36 | 0.45 |
| sqlite-graph-memory (graph) | 0.85 | 0.36 | 0.45 |
| BM25, no model | 0.77 | 0.45 | 0.36 |
On "what is the current deadline" questions, retrieval is perfect (R@5 1.00). The system still quotes the superseded decision 6 times out of 6.
CI recomputes every number from the per-question outputs. It reran all three rows from a clean index on Linux and matched exactly: drift 0.0000, 0 of 91 rankings changed. A forged result is rejected (#20).
Add your system: bench/CONTRIBUTING.md. Closes #16.
Built by AI agents (Mycroft, Anton's synthetic AI cofounder, with Codex reviewing). Maintained by a human.
v0.2.1 - CI that proves the red before it trusts the green
The suite stopped being something a human remembers to run. The 84 tests and the eval suite shipped in v0.2.0 had only ever been run by hand. .github/workflows/tests.yml now runs them on every push against Python 3.11, 3.12 and 3.13.
It carries one step that is not boilerplate: the workflow sets EVAL_MUTANT=1, which breaks the retrieval pipeline on purpose, and fails the build if the eval suite still comes back green. A test that has never been seen red is not evidence of anything, so the pipeline proves its own detection before it trusts a passing run. If someone later guts the retrieval stages and the eval keeps smiling, CI stops the push instead of the user discovering it.
Read this with AI. One click opens the repo in Codex, ChatGPT or Claude with a prompt that asks the agent to extract the reusable patterns and apply them to your own setup; the raw prompt is there for any other model. The README also names the neighbouring repos in the memory layer, so the piece can be placed in the system it was lifted out of.
One correction. A commit on 30 September "corrected" a measurement date in the docs, and it was reverted the same day: the original date was right, and the machine making the correction was the one with the wrong clock. Recorded rather than quietly dropped, because the repo's own argument is that the measurement date matters.
v0.2.0 — an eval that says no, and an MCP server that says nothing correctly
Two halves shipped since v0.1.3: a retrieval eval that exists to talk us out of upgrades, and an MCP server you can actually point a harness at.
The eval, and the seven upgrades it talked us out of. eval/build_gold.py turns a vault's own [[wikilinks]] into a retrieval test set — no hand labelling, no LLM calls — and eval/run_eval.py scores it as Recall@12 / MRR / nDCG@12 in both modes, vector-only and vector+graph. It imports the pipeline rather than copying it, so what it measures is the thing that ships. Four question classes, because "does the graph help?" has four different answers: title, body, bridge, temporal.
The README now carries real numbers from the home vault: two changes that improved things (bridge Recall@12 0.251 → 0.451) and seven that obviously should have helped and measurably did not — BGE-M3, two rerankers, a BM25 hybrid, a heading-path chunker, sqlite-vec, and more. That list is the point of the module.
The ruler was broken, and an outside panel found it first. Before announcing the eval anywhere it went to two independent review engines, and both landed on the same defect: ndcg_at credited the same gold note once per appearance in the ranking, so ndcg_at(['foo','foo'], ['foo']) returned 1.63. Wikilinks address notes by filename, so two folders holding a foo.md is not a corner case. Fixed, with the published numbers corrected downward (title nDCG 0.921 → 0.890, temporal 0.638 → 0.531). The deltas survived to the fourth decimal, because the same four questions inflated both the before and the after run — so the merge/revert decisions taken from those runs still stand. EVAL_MUTANT=1 breaks nDCG on purpose and the suite must go red under it: a test that stays green while the metric is broken never tested the metric.
The MCP server got a runtime and three fixes to how it reports nothing. examples/ now carries a runnable server and its own README (#5). The answer file is scoped per call, so two recalls in flight can no longer return each other's text. An empty answer file used to fall back to whatever was on stdout, which meant a log line could come back as a recall result — it returns isError now (#9) — and a file it could not read is reported separately from a recall that genuinely found nothing (#13).
The paper is published: JOSS draft plus CITATION.cff, DOI 10.5281/zenodo.22639718.
Full journal: CHANGELOG.md
Built in the open by Anton Dzyatkovskiy and Mycroft, his synthetic co-founder. More of the same at https://github.com/tonydzi — if you run agents with long memory and want a second pair of eyes on this, the issues are open.
v0.1.3 — not everything ending in .md is a note
Two ways the indexer quietly answered with the wrong text. Both were found on a real synced vault rather than in review, and both now have tests that run with no model download and no network.
Not everything ending in .md is a note. The walk was a plain rglob('*.md'), which descends into dot-directories. On an Obsidian vault synced with Syncthing that pulled in .stversions/ (Syncthing's old revisions), *.sync-conflict-* copies, .obsidian/ plugin cache and .git/ internals.
This is worse than ordinary noise. A stale revision of a note is semantically almost identical to the live one, so it does not sit harmlessly at the bottom of the ranking — it lands right next to the real note in the top-K and answers the question with outdated content. The failure is silent by construction: retrieval looks healthy, the answer is just wrong.
Those paths are now skipped. BRAIN_INDEX_HIDDEN=1 opts back in, because someone will legitimately keep notes in a hidden folder.
A byte-order mark ate the frontmatter. Notes were read as plain utf-8, so a BOM'd file kept a leading U+FEFF and its first line was no longer ---. The frontmatter regex then saw a note with no frontmatter and dropped its date and title without complaint. Reading as utf-8-sig fixes it.
The two behaviours are now iter_notes and read_note, which is what made them testable at all. tests/test_index_notes.py has seven cases over a directory shaped like a real synced vault; they were checked red against the pre-fix behaviour before being called done.
Closes #1 and #2 — #2 asked for a smoke test that runs without downloading a model, and this is one.
v0.1.2 — the bilingual regex, explained
docs patch:
- why the decision regex is bilingual, written down next to the regex
- dead org link fixed after the account rename; contact footer with the engineer CTA
cut by the weekly release pass, first run.
v0.1.1 — the design note, and everything a stranger needs
Everything the pilot gained after the first tag: the design note that explains where it is going, and the parts a stranger needs in order to use or contribute to it.
New
docs/bitemporal.md— the design note for bi-temporal edges. When you materialize the graph instead of parsing it at query time, give every edge a validity window (valid_from/valid_to/observed_at) so a rebuild closes superseded facts instead of deleting them: history is kept, recall prefers the present,as_ofqueries can reconstruct the past. The day-1 A/B result — context size ~unchanged, candidates ~60% fresher by mean age — is written up with its caveat: one day, one corpus, longitudinal effect, not a benchmark.- A roadmap in the README with Now/Next, and the two known defects named as open issues rather than footnotes.
AGENTS.md— for the reader who is an agent: it is a pilot, and the absent parts are deliberate.FOR-ROBOTS.md,CITATION.cff, per-claim attribution in the README (which claim is backed by what), the contributor deal (no CLA, you keep your copyright, 48h answer), and changelog categories for these auto-generated notes.
Known defects, still open
- #1 — the indexer ingests
.stversionsbackups, sync-conflict copies and.obsidianjunk as if they were notes: five files where there is one. - #2 — no tests at all, on 677 lines.
Both are scoped and free to take. Until #2 exists, pilot is the only word this repo is entitled to, and the README says so.
What's next
A smoke test that runs without downloading a model, ignore rules for the indexer, then v0.2 — a public benchmark with a synthetic mini-vault, ~200 hand-labeled queries stratified by type, and a full ablation matrix. The interesting question is not "does graph help" but for which query classes. That line is a plan, not a result.
From here on every noticeable change ships as its own release, so this feed — not the commit graph — is where you can see whether "pilot" has stopped being the right word.
Full Changelog: v0.1.0...v0.1.1
v0.1.0 — the pilot as first published (3 July 2026)
Graph RAG on SQLite for AI agents — a working pilot, not a framework.
Backfilled release note. The
v0.1.0tag has pointed at this commit since 3 July 2026; what was missing was the changelog, not the code. It is written up now so the release feed tells the truth about when each state of this repo actually existed.
What v0.1.0 is
The extracted memory layer of a personal second-brain agent setup — three small Python scripts that give an LLM agent associative recall over a folder of markdown notes. SQLite is the only database and [[wikilinks]] are the graph.
index_notes.py— chunk + embed (e5-base) into a.npy/.pklindex.brain_ask.py— the recall pipeline in one file, in order: dense retrieve → optional 1-hop wikilink expansion → cross-encoder rerank →--abmode that runs vector-only and vector+graph and logs the delta to SQLite.turnstate_hook.py— a Stop hook that appends one row per assistant turn (ask, summary, files, tools, commands, decisions). Zero LLM tokens, pure stdlib.turnstate_show.pyreads it back.schema.sql— both tables, documented.
The bet it encodes
Most Graph RAG stacks assume a graph database, an ETL pipeline and an entity-extraction pass. For a single-user agent over a markdown knowledge base, all three are overkill: the graph already exists because the human hand-curated it as wikilinks, SQLite is enough for the only things worth persisting, and the expensive part of RAG quality is the reranker, not graph infrastructure.
Two design choices that survived contact with reality are in this release: graph expansion is candidate generation and not ranking (so an irrelevant linked note gets buried by the reranker), and an entity gate that switches the hop off for name-shaped queries — because a person's card links to everything, and A/B telemetry showed the hop helping theme queries while hurting name lookups.
What it is not
Pilot is load-bearing: no tests, no eval suite, no incremental indexing, no packaging, no entity lane. It runs daily on one real ~100k-note vault. That is a use, not a benchmark.
Full Changelog: https://github.com/Palo-Alto-AI-Research-Lab/sqlite-graph-memory/commits/v0.1.0