Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
226 changes: 82 additions & 144 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,23 +1,32 @@
# Memrank

Memrank is a tool for reproducible, auditable evaluation of memory systems.
[![PyPI release](https://img.shields.io/pypi/v/memrank)](https://pypi.org/project/memrank/)
[![Code license: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue)](https://github.com/atomicstrata/memrank/blob/main/LICENSE)

**Status:** v0.4, in active development. Interfaces still move between releases.

## Quick start
Memrank is a tool for reproducible, auditable evaluation of memory systems.

Paste this to your agent, and it installs memrank and runs the evaluation below for you:
Compare memory systems and their versions on tasks that matter to your agent or application.
You choose the tasks and success criteria; Memrank runs the evaluation, reports scores and timings,
and records responses and errors so you can investigate differences.

```text
Install memrank in this project and run its smoke evaluation, following
https://github.com/atomicstrata/memrank/blob/main/docs/install.md. Check the prerequisites
that page lists before you change anything, install into this project only, and do not
install anything globally or edit my shell configuration. When the run finishes, show me the
`system:` and `evaluation:` lines it printed. Stop and ask me if any step fails.
```
[Start here](https://github.com/atomicstrata/memrank/blob/main/docs/getting-started.md) |
[Compare memory systems and their versions](https://github.com/atomicstrata/memrank/blob/main/docs/comparing.md) |
[Understand results](https://github.com/atomicstrata/memrank/blob/main/docs/results.md)

## Quick start

Check the installation with a small local retrieval evaluation. Installation needs the network;
the Python block below needs no key, service or network. This checks that Memrank works in your
project; it does not establish how a memory system will perform on your workload.

### Do it yourself

Use Python >= 3.10 and an existing project or virtualenv.
[Installation instructions](https://github.com/atomicstrata/memrank/blob/main/docs/install.md)
cover creating one.

```bash
uv add memrank # or, into a virtualenv you already have: pip install memrank
```
Expand All @@ -33,172 +42,101 @@ print(result)
```

[`TFIDF`](https://github.com/atomicstrata/memrank/blob/main/docs/systems/tfidf.md) is keyword search
weighted by how rare each word is, and it ships with the package.
weighted by how rare each word is.
[`SQuAD`](https://github.com/atomicstrata/memrank/blob/main/docs/evaluations/squad.md) supplies
32 bundled passages and 64 questions. This measures full-passage retrieval recall, not
answer-span or end-to-end answer correctness. Neither needs an
[engine](https://github.com/atomicstrata/memrank/blob/main/docs/reference/system.md#system-and-engine),
a key or the network.
[Installing memrank](https://github.com/atomicstrata/memrank/blob/main/docs/install.md) covers uv,
Python versions and upgrading.
32 bundled passages and 64 questions. The `squad-score` measures full-passage retrieval recall,
not answer-span or end-to-end answer correctness. Check that the output names the expected system
and evaluation and says `64 recorded, 0 with errors`.
[Read the output](https://github.com/atomicstrata/memrank/blob/main/docs/getting-started.md#read-the-installation-check).

## Use cases
Or paste this to your coding agent:

Each snippet below runs on its own.

### Evaluate a system of your own

Four methods, and the
[system](https://github.com/atomicstrata/memrank/blob/main/docs/reference/system.md) is ready to
evaluate. A memory engine you already run has a client that ships with memrank instead --
`AtomicMemory`, `Hindsight`, `Mem0` and `Supermemory` take a `base_url=` where `NoteBook()` goes
below, and [adding a system](https://github.com/atomicstrata/memrank/blob/main/docs/systems.md) is
the rest.

```python
from memrank import Memory, Recall
from memrank.evaluations import Demo


class NoteBook(Memory):
name, version, engine_version = "notebook", "0.1", "0.1"

def prepare(self, isolation_unit):
self.notes = []
```text
Install memrank in this project and run its smoke evaluation, following
https://github.com/atomicstrata/memrank/blob/main/docs/install.md. Check the prerequisites
that page lists before you change anything, install into this project only, and do not
install anything globally or edit my shell configuration. When the run finishes, show me the
`system:`, `evaluation:` and `traces:` lines it printed. Stop and ask me if any step fails.
```

def ingest(self, documents):
self.notes.extend(documents)
## Use cases

def retrieve(self, query, k, user_id, query_timestamp=None) -> Recall:
wanted = set(query.lower().split())
ranked = sorted(self.notes, reverse=True,
key=lambda note: len(wanted & set(note.content.lower().split())))
return Recall(documents=ranked[:k])
<a id="compare-two-systems"></a>
### Compare two memory systems

def cleanup(self):
self.notes = []
Run candidates on the same evaluation, inspect coverage and failures, and read each measure's
meaning before interpreting a gap. The [memory systems comparison guide](https://github.com/atomicstrata/memrank/blob/main/docs/comparing.md)
shows how to read the differences, with an offline toy example of task-level pairing.
It also explains when to compare summaries: `memrank.paired` does not compare aggregate
scores such as `squad-score`.

<a id="evaluate-a-system-of-your-own"></a>
### Evaluate a memory system of your own

result = Demo().run(system=NoteBook())
```
Use a [shipped client](https://github.com/atomicstrata/memrank/blob/main/docs/systems/README.md)
or [connect your own system](https://github.com/atomicstrata/memrank/blob/main/docs/systems.md).
A memory implements `prepare`, `ingest`, `retrieve` and `cleanup` so Memrank can give it context,
ask questions and clear state between independent cases. An external
[engine](https://github.com/atomicstrata/memrank/blob/main/docs/reference/system.md#system-and-engine)
may need a configured service and credentials.

### Ask your own questions

An [evaluation](https://github.com/atomicstrata/memrank/blob/main/docs/reference/evaluation.md) you
write by hand and one that ships are the same object:
[tasks](https://github.com/atomicstrata/memrank/blob/main/docs/reference/task.md), the
[measures](https://github.com/atomicstrata/memrank/blob/main/docs/measures.md) that read them, and
when the system is cleared.

```python
from memrank import Clearing, Document, Evaluation, Expected, Task, WordMatch
from memrank.systems import WordOverlap

notes = (Document(id="t1", user_id="acme",
content="Acme moved to the enterprise plan in March."),)

tickets = Evaluation(
name="tickets", version="internal@2026-09",
tasks=(Task(id="q_plan", prompt="What plan is Acme on?", group="acme", context=notes,
expected=Expected(required_spans=("enterprise",), evidence_doc_ids=("t1",))),),
measures=(WordMatch(),), clearing=Clearing.PER_GROUP)

result = tickets.run(system=WordOverlap())
```
[Express your evaluation](https://github.com/atomicstrata/memrank/blob/main/docs/evaluations.md)
as tasks, expected outcomes and measures. You decide which cases represent your problem and what
counts as success; Memrank applies those rules and records the evidence.

### Find out why a value is what it is

Every [value](https://github.com/atomicstrata/memrank/blob/main/docs/reference/result.md) names the
task it came from, and every task kept its
[trace](https://github.com/atomicstrata/memrank/blob/main/docs/reference/trace.md) -- what was
asked, what the evaluation wanted, what came back.

```python
from memrank.evaluations import Demo
from memrank.systems import WordOverlap

result = Demo().run(system=WordOverlap())

lowest = min(result.values_of("word-match"), key=lambda value: value.value or 0.0)
trace = result.traces_of(lowest.task_id)[0]

print(lowest.value, lowest.why)
print(trace.task.prompt, trace.task.expected.required_spans)
for document in trace.recalled.documents:
print(document.id, document.content)
```

### Compare two systems

[`memrank.paired`](https://github.com/atomicstrata/memrank/blob/main/docs/reference/paired.md)
refuses two results of different evaluations, then reads them task by task and says how often chance
alone produces a gap that size. It never says "better".

```python
import memrank
from memrank.evaluations import Demo
from memrank.systems import NoContext, WordOverlap

evaluation = Demo()

print(memrank.paired(evaluation.run(system=WordOverlap()),
evaluation.run(system=NoContext())))
```
[Understand results](https://github.com/atomicstrata/memrank/blob/main/docs/results.md) shows how
to inspect a task's trace, distinguish missing values from zero, and save a result for later use.
[Write a measure](https://github.com/atomicstrata/memrank/blob/main/docs/measures.md) to read
something new from stored traces without rerunning the system.

### Check the instrument

`NoContext` is the floor: it retrieves nothing. `FullContext` is the ceiling: it is given every
document, unranked. A gap between them is what makes the evaluation worth running at all, and
[methodology](https://github.com/atomicstrata/memrank/blob/main/docs/methodology.md) states what a
value does and does not license you to say.

```python
from memrank.evaluations import Demo
from memrank.systems import FullContext, NoContext, WordOverlap

evaluation = Demo()

for system in (NoContext(), WordOverlap(), FullContext()):
result = evaluation.run(system=system)
scored = [value.value for value in result.values_of("word-match")
if value.value is not None]
print(result.system.name, sum(scored) / len(scored))
```
[Control examples](https://github.com/atomicstrata/memrank/blob/main/examples/05-against-a-baseline/README.md)
show what happens when retrieval returns no documents or all documents. These checks help expose
what a measure rewards; they do not prove that the evaluation represents your workload.
[Methodology](https://github.com/atomicstrata/memrank/blob/main/docs/methodology.md) states the
measurement rules and limits on claims.

## Where to read more

| | |
| Guide | Task |
|---|---|
| [Systems](https://github.com/atomicstrata/memrank/blob/main/docs/systems/README.md) | one page per system memrank ships, and what each one needs |
| [Evaluations](https://github.com/atomicstrata/memrank/blob/main/docs/evaluations/README.md) | one page per evaluation, its tasks and what it measures |
| [Reference](https://github.com/atomicstrata/memrank/blob/main/docs/reference/README.md) | one page per word in the Python surface |
| [Measures](https://github.com/atomicstrata/memrank/blob/main/docs/measures.md) | what a scoring rule declares, and the ones memrank ships |
| [Methodology](https://github.com/atomicstrata/memrank/blob/main/docs/methodology.md) | the axes, the budget control, the control arms, evidence classes |
| [Installing memrank](https://github.com/atomicstrata/memrank/blob/main/docs/install.md) | prerequisites, install, a smoke run, upgrading |
| [Local development](https://github.com/atomicstrata/memrank/blob/main/docs/local-development.md) | working on memrank itself |
| [SPEC.md](https://github.com/atomicstrata/memrank/blob/main/docs/SPEC.md) | what memrank evaluates, and the governance it commits to |

Memrank ships a `memrank` command as well, and it is not core: nothing above needs it, and it keeps
an older vocabulary of its own --
[the command line](https://github.com/atomicstrata/memrank/blob/main/docs/misc/command-line.md) is
where it lives.
| [Start here](https://github.com/atomicstrata/memrank/blob/main/docs/getting-started.md) | Start using Memrank |
| [Memory systems comparison guide](https://github.com/atomicstrata/memrank/blob/main/docs/comparing.md) | Compare memory systems or their versions |
| [Understand results](https://github.com/atomicstrata/memrank/blob/main/docs/results.md) | Interpret and save measurements |
| [Systems](https://github.com/atomicstrata/memrank/blob/main/docs/systems/README.md) | Find available integrations |
| [Evaluations](https://github.com/atomicstrata/memrank/blob/main/docs/evaluations/README.md) | Find available task sets |
| [Reference](https://github.com/atomicstrata/memrank/blob/main/docs/reference/README.md) | Look up Python contracts |
| [Documentation index](https://github.com/atomicstrata/memrank/blob/main/docs/README.md) | Find every guide |

The guides above use Python. Memrank also provides a
[command line](https://github.com/atomicstrata/memrank/blob/main/docs/misc/command-line.md) for
tracked and placed runs, and a
[translator contract](https://github.com/atomicstrata/memrank/blob/main/docs/system-contract.md)
for memory systems implemented in other languages.

## Help and contribution

[Contributing and getting help](https://github.com/atomicstrata/memrank/blob/main/docs/contributing.md)
links the issue tracker and development checks. Questions about your setup are easier to reproduce
with the package version, a small example and the error text. Do not include keys or private data.

## Governance

Memrank is maintained by [AtomicStrata](https://atomicstrata.ai) under a vendor-neutral charter:
anyone may submit a system, results are published as measured, and methodology changes go through
public proposal and comment. The commitments are in
[SPEC.md section 7](https://github.com/atomicstrata/memrank/blob/main/docs/SPEC.md#7-governance----the-vendor-neutral-charter).
AtomicStrata
also ships a memory engine, AtomicMemory, which this tool evaluates and which has placed below a
no-memory control arm in our own runs -- which is why the floor and the ceiling above are in the
package rather than in a report of ours.
AtomicStrata also develops AtomicMemory, one of the engines Memrank can evaluate. Comparisons
should be assessed through their method, configuration and recorded evidence.

## Licences

Memrank's code is [Apache-2.0](https://github.com/atomicstrata/memrank/blob/main/LICENSE).
The bundled SQuAD subset is [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/);
its [notice](https://github.com/atomicstrata/memrank/blob/main/memrank/benchmarks/data/SQUAD-NOTICE.md)
credits the creators and passage sources and records the selection and reformatting.

Methodology questions and disagreements: open an issue. Anything else: hello@atomicstrata.ai
Loading
Loading