A pluggable memory engine for AI agents. Five pieces, BYO database / LLM / embedder.
1. Hybrid retrieval (pgvector + FTS + RRF, agentic RAG via LLM tools)
2. Agent-Being doc (worn first-person self-model, loaded every prompt)
3. Anticipation bundles (event-triggered, pre-warmed for O(1) reply-time reads)
4. Sleep-time writer (nightly LLM rewrite of the doc from the day's events)
5. Foreground draft loop (the integration point in your handler)
The agent has a first-person document about its user, regenerated nightly from the day's events, loaded into every prompt unconditionally. Retrieval is the fallback. Anticipation is event-driven. Sleep is when integration happens.
This is consensus 2025–26 architecture (Letta sleep-time, Anthropic Auto Dream, Mastra Observer/Reflector, Hindsight). The novelty isn't the architecture — it's that the worn doc is the front of memory, not a derivative summary. Most "memory systems" make retrieval the front; here, retrieval is the fallback.
flowchart TB
subgraph Host["Host application"]
Webhook["Inbound webhook<br/>(email · WhatsApp · webhook)"]
Reply["Reply handler<br/>(draft + send)"]
Cron["Daily cron<br/>(4am IST)"]
end
subgraph WM["worn-memory engine"]
Being[("Agent-Being doc<br/>(worn · every prompt)")]
Anticipation[("Anticipation bundles<br/>(event-keyed · TTL'd)")]
EventLog[("Event log<br/>(append-only)")]
Tools["recall_memory<br/>recent_memory<br/>(LLM tools)"]
Dream["Sleep-time writer"]
end
subgraph BYO["Bring your own"]
DB[("Postgres + pgvector")]
LLM["LLM provider"]
Embed["Embedder"]
end
Webhook -->|persist| EventLog
Webhook -->|ensureBeing| Being
Webhook -.->|fire-and-forget| Anticipation
Reply -->|read worn doc| Being
Reply -.->|optional fallback| Tools
Tools -->|retrieve| EventLog
Cron -->|dreamAll| Dream
Dream -->|read 24h| EventLog
Dream -->|read current| Being
Dream -->|rewrite| Being
Being -.uses.-> DB
Anticipation -.uses.-> DB
EventLog -.uses.-> DB
Tools -.embed.-> Embed
Tools -.search.-> DB
Anticipation -.LLM call.-> LLM
Dream -.LLM call.-> LLM
Solid arrows are synchronous calls in the host's flow. Dotted arrows are fire-and-forget or background operations. The dashed lines into the BYO column are dependency edges — things you provide via adapters.
For the write-path and sleep-cycle sequence diagrams, see ARCHITECTURE.md.
This is source-distributed. Vendor it into your project (Git submodule, copy paste, or fork). An npm package release is on the roadmap; until then, drop src/ into lib/memory/ of your repo.
git clone https://github.com/balabommablock-cpu/worn-memory.git lib/memoryPeer dependencies your project needs:
| dep | why |
|---|---|
drizzle-orm |
for the reference Postgres adapter |
ai (Vercel AI SDK) |
for the LLM tool wrappers (recall_memory, recent_memory) |
zod |
for tool parameter schemas |
Postgres + pgvector |
for the retrieval substrate |
If you aren't on Drizzle / Vercel AI SDK, see INTEGRATION.md — the adapter contracts in src/types.ts are stack-agnostic.
import { createMemory } from 'worn-memory';
import { drizzleAdapters } from 'worn-memory/adapters/drizzle';
import { db } from './db';
import { agents, users, memories } from './schema';
import { agentBeing, anticipations } from './schema/memory'; // copy from worn-memory/schema
const memory = createMemory({
...drizzleAdapters({
db,
tables: { agentBeing, anticipations, agents, memories },
events: {
// What the dream cron reads when regenerating the worn doc.
// Implement against whatever you consider a "meaningful event":
// inbound messages, completed actions, user overrides, etc.
async recent({ agentId, since, limit }) { /* ... */ },
},
}),
llm: { generate: async (args) => myLLM.call(args) },
embed: { embed: async (text) => myEmbedder(text) },
});Then:
// In your inbound webhook
const being = await memory.ensureBeing(agent);
void memory.runAnticipation({ trigger: 'email.inbound', triggerRefId, snippet });
const reply = await llm.generate({
system: `You are ${agent.name}.\n\n${being}\n\n...`,
tools: memory.tools(agent.id),
});
// In a nightly cron
await memory.dreamAll();worn-memory/
├── README.md ← you are here
├── ARCHITECTURE.md ← the five pieces, why each is load-bearing
├── INTEGRATION.md ← adapter contracts; how to plug into non-Drizzle stacks
├── LICENSE ← MIT
├── src/
│ ├── index.ts ← public API: createMemory({ adapters }) → Memory
│ ├── types.ts ← adapter contracts (Search, Being, Anticipation, EventSource, AgentRegistry, LLM, Embed)
│ ├── primitive.ts ← hybrid retrieval, RRF-fused
│ ├── tool.ts ← LLM tools wrapping the primitive
│ ├── agent-being.ts ← worn doc — getBeing / ensureBeing / writeBeing
│ ├── anticipate.ts ← event-driven runAnticipation + latestAnticipation
│ ├── dream.ts ← sleep-time writer — dreamForAgent / dreamAll
│ ├── _id.ts ← minimal ulid (overridable)
│ └── adapters/
│ └── drizzle.ts ← reference Postgres + Drizzle adapter
├── schema/
│ └── drizzle.ts ← Drizzle schema for `agent_being` + `anticipations`
├── migrations/
│ └── 0001_memory_engine.sql ← portable Postgres DDL
└── examples/
├── basic-usage.ts ← minimum wiring
└── nextjs-cron.route.ts ← reference cron route
const memory = createMemory(config);
// Retrieval (agentic RAG)
memory.search(agentId, query, k?) → MemoryHit[]
memory.recent(agentId, n?) → MemoryHit[]
memory.tools(agentId) → { recall_memory, recent_memory } // AI SDK shape
// Worn doc
memory.getBeing(agentId) → string
memory.ensureBeing(agent) → string (seeds on first call)
memory.writeBeing(agentId, doc) → void
// Event-driven anticipation
memory.runAnticipation({ trigger, ... }) → string | null (bundle id)
memory.latestAnticipation({ trigger, ... }) → AnticipationBundle | null
memory.consumeAnticipation(id) → void
// Sleep-time writer
memory.dreamForAgent(agentId) → DreamResult
memory.dreamAll() → { total, updated, skipped, errored, results }See src/index.ts and src/types.ts for full signatures.
Most memory systems retrieve. You ask, they fetch. That's a good answer for "find me a doc," and a thin one for "be the same agent over time." For a personal companion AI, identity-memory matters as much as retrieval-memory — and it has a different structural shape.
The worn doc is the agent's self-model — what it understands about you, how it intends to act, what's currently on its mind. It's loaded into every prompt unconditionally. The agent never retrieves its self-model; it wears it. Retrieval handles the cases the doc doesn't (specific facts, deep history, anomalies), but the doc handles the 80% of "be present, know me" interactions for free.
Anticipation pre-warms the next move. When something happens — an email, a message, a webhook — the agent doesn't wait until the user asks before reasoning. A focused LLM pass fires immediately, stashes a bundle. By the time the user pings about it, the answer is already warm. Event-driven, not interval cron.
The dream cycle is when the agent integrates. Once a day, a writer agent reads the day's events and rewrites the worn doc. Continuity is preserved (yesterday's truths stay unless explicitly contradicted); new learning is integrated; stale concerns decay. The user feels the same agent, but slightly more knowing of them, every morning.
That's it. No bitemporal graphs (yet). No belief networks (yet). No fine-tuning per user. Just plain text, plain tables, plain LLM calls. The substrate is boring. The loop is the product.
These are forward-compatible additions, not breaking changes:
- Bitemporal fact graph — Graphiti-style, with
valid_from/valid_toper edge. Useful for "what did the user believe about X in March?" queries that the worn doc cannot answer precisely. - Hindsight-style belief network — separate stores for facts vs experiences vs entity summaries vs evolving beliefs.
- A-Mem self-evolving links — when a new memory is written, retroactively update old memories' tags/links.
- Adaptive retrieval (RMM-style) — track which retrievals were useful and weight future calls accordingly.
- Daily consolidation patterns — extract recurring patterns from the event log and promote them as first-class memories.
The current architecture is sized for a 100-to-100k-user consumer agent. The above are sprint-2 candidates if usage justifies them.
MIT. See LICENSE.