Skip to content

About

Pluggable memory engine for AI agents — worn first-person doc + event-driven anticipation + sleep-time writer + hybrid retrieval. BYO database / LLM / embedder.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

worn-memory

A pluggable memory engine for AI agents. Five pieces, BYO database / LLM / embedder.

1. Hybrid retrieval        (pgvector + FTS + RRF, agentic RAG via LLM tools)
2. Agent-Being doc         (worn first-person self-model, loaded every prompt)
3. Anticipation bundles    (event-triggered, pre-warmed for O(1) reply-time reads)
4. Sleep-time writer       (nightly LLM rewrite of the doc from the day's events)
5. Foreground draft loop   (the integration point in your handler)

The agent has a first-person document about its user, regenerated nightly from the day's events, loaded into every prompt unconditionally. Retrieval is the fallback. Anticipation is event-driven. Sleep is when integration happens.

This is consensus 2025–26 architecture (Letta sleep-time, Anthropic Auto Dream, Mastra Observer/Reflector, Hindsight). The novelty isn't the architecture — it's that the worn doc is the front of memory, not a derivative summary. Most "memory systems" make retrieval the front; here, retrieval is the fallback.


architecture at a glance

flowchart TB
    subgraph Host["Host application"]
        Webhook["Inbound webhook<br/>(email · WhatsApp · webhook)"]
        Reply["Reply handler<br/>(draft + send)"]
        Cron["Daily cron<br/>(4am IST)"]
    end

    subgraph WM["worn-memory engine"]
        Being[("Agent-Being doc<br/>(worn · every prompt)")]
        Anticipation[("Anticipation bundles<br/>(event-keyed · TTL'd)")]
        EventLog[("Event log<br/>(append-only)")]
        Tools["recall_memory<br/>recent_memory<br/>(LLM tools)"]
        Dream["Sleep-time writer"]
    end

    subgraph BYO["Bring your own"]
        DB[("Postgres + pgvector")]
        LLM["LLM provider"]
        Embed["Embedder"]
    end

    Webhook -->|persist| EventLog
    Webhook -->|ensureBeing| Being
    Webhook -.->|fire-and-forget| Anticipation
    Reply -->|read worn doc| Being
    Reply -.->|optional fallback| Tools
    Tools -->|retrieve| EventLog
    Cron -->|dreamAll| Dream
    Dream -->|read 24h| EventLog
    Dream -->|read current| Being
    Dream -->|rewrite| Being

    Being -.uses.-> DB
    Anticipation -.uses.-> DB
    EventLog -.uses.-> DB
    Tools -.embed.-> Embed
    Tools -.search.-> DB
    Anticipation -.LLM call.-> LLM
    Dream -.LLM call.-> LLM
Loading

Solid arrows are synchronous calls in the host's flow. Dotted arrows are fire-and-forget or background operations. The dashed lines into the BYO column are dependency edges — things you provide via adapters.

For the write-path and sleep-cycle sequence diagrams, see ARCHITECTURE.md.


install

This is source-distributed. Vendor it into your project (Git submodule, copy paste, or fork). An npm package release is on the roadmap; until then, drop src/ into lib/memory/ of your repo.

git clone https://github.com/balabommablock-cpu/worn-memory.git lib/memory

Peer dependencies your project needs:

dep why
drizzle-orm for the reference Postgres adapter
ai (Vercel AI SDK) for the LLM tool wrappers (recall_memory, recent_memory)
zod for tool parameter schemas
Postgres + pgvector for the retrieval substrate

If you aren't on Drizzle / Vercel AI SDK, see INTEGRATION.md — the adapter contracts in src/types.ts are stack-agnostic.


quick start (Postgres + Drizzle)

import { createMemory } from 'worn-memory';
import { drizzleAdapters } from 'worn-memory/adapters/drizzle';

import { db } from './db';
import { agents, users, memories } from './schema';
import { agentBeing, anticipations } from './schema/memory';   // copy from worn-memory/schema

const memory = createMemory({
  ...drizzleAdapters({
    db,
    tables: { agentBeing, anticipations, agents, memories },
    events: {
      // What the dream cron reads when regenerating the worn doc.
      // Implement against whatever you consider a "meaningful event":
      // inbound messages, completed actions, user overrides, etc.
      async recent({ agentId, since, limit }) { /* ... */ },
    },
  }),
  llm:   { generate: async (args) => myLLM.call(args) },
  embed: { embed:    async (text) => myEmbedder(text) },
});

Then:

// In your inbound webhook
const being = await memory.ensureBeing(agent);
void memory.runAnticipation({ trigger: 'email.inbound', triggerRefId, snippet });

const reply = await llm.generate({
  system: `You are ${agent.name}.\n\n${being}\n\n...`,
  tools: memory.tools(agent.id),
});

// In a nightly cron
await memory.dreamAll();

what's in the box

worn-memory/
├── README.md              ← you are here
├── ARCHITECTURE.md        ← the five pieces, why each is load-bearing
├── INTEGRATION.md         ← adapter contracts; how to plug into non-Drizzle stacks
├── LICENSE                ← MIT
├── src/
│   ├── index.ts           ← public API: createMemory({ adapters }) → Memory
│   ├── types.ts           ← adapter contracts (Search, Being, Anticipation, EventSource, AgentRegistry, LLM, Embed)
│   ├── primitive.ts       ← hybrid retrieval, RRF-fused
│   ├── tool.ts            ← LLM tools wrapping the primitive
│   ├── agent-being.ts     ← worn doc — getBeing / ensureBeing / writeBeing
│   ├── anticipate.ts      ← event-driven runAnticipation + latestAnticipation
│   ├── dream.ts           ← sleep-time writer — dreamForAgent / dreamAll
│   ├── _id.ts             ← minimal ulid (overridable)
│   └── adapters/
│       └── drizzle.ts     ← reference Postgres + Drizzle adapter
├── schema/
│   └── drizzle.ts         ← Drizzle schema for `agent_being` + `anticipations`
├── migrations/
│   └── 0001_memory_engine.sql  ← portable Postgres DDL
└── examples/
    ├── basic-usage.ts            ← minimum wiring
    └── nextjs-cron.route.ts      ← reference cron route

the public API

const memory = createMemory(config);

// Retrieval (agentic RAG)
memory.search(agentId, query, k?)             → MemoryHit[]
memory.recent(agentId, n?)                    → MemoryHit[]
memory.tools(agentId)                         → { recall_memory, recent_memory }   // AI SDK shape

// Worn doc
memory.getBeing(agentId)                      → string
memory.ensureBeing(agent)                     → string  (seeds on first call)
memory.writeBeing(agentId, doc)               → void

// Event-driven anticipation
memory.runAnticipation({ trigger, ... })      → string | null   (bundle id)
memory.latestAnticipation({ trigger, ... })   → AnticipationBundle | null
memory.consumeAnticipation(id)                → void

// Sleep-time writer
memory.dreamForAgent(agentId)                 → DreamResult
memory.dreamAll()                             → { total, updated, skipped, errored, results }

See src/index.ts and src/types.ts for full signatures.


why this shape

Most memory systems retrieve. You ask, they fetch. That's a good answer for "find me a doc," and a thin one for "be the same agent over time." For a personal companion AI, identity-memory matters as much as retrieval-memory — and it has a different structural shape.

The worn doc is the agent's self-model — what it understands about you, how it intends to act, what's currently on its mind. It's loaded into every prompt unconditionally. The agent never retrieves its self-model; it wears it. Retrieval handles the cases the doc doesn't (specific facts, deep history, anomalies), but the doc handles the 80% of "be present, know me" interactions for free.

Anticipation pre-warms the next move. When something happens — an email, a message, a webhook — the agent doesn't wait until the user asks before reasoning. A focused LLM pass fires immediately, stashes a bundle. By the time the user pings about it, the answer is already warm. Event-driven, not interval cron.

The dream cycle is when the agent integrates. Once a day, a writer agent reads the day's events and rewrites the worn doc. Continuity is preserved (yesterday's truths stay unless explicitly contradicted); new learning is integrated; stale concerns decay. The user feels the same agent, but slightly more knowing of them, every morning.

That's it. No bitemporal graphs (yet). No belief networks (yet). No fine-tuning per user. Just plain text, plain tables, plain LLM calls. The substrate is boring. The loop is the product.


what's not built (yet)

These are forward-compatible additions, not breaking changes:

  • Bitemporal fact graph — Graphiti-style, with valid_from / valid_to per edge. Useful for "what did the user believe about X in March?" queries that the worn doc cannot answer precisely.
  • Hindsight-style belief network — separate stores for facts vs experiences vs entity summaries vs evolving beliefs.
  • A-Mem self-evolving links — when a new memory is written, retroactively update old memories' tags/links.
  • Adaptive retrieval (RMM-style) — track which retrievals were useful and weight future calls accordingly.
  • Daily consolidation patterns — extract recurring patterns from the event log and promote them as first-class memories.

The current architecture is sized for a 100-to-100k-user consumer agent. The above are sprint-2 candidates if usage justifies them.


license

MIT. See LICENSE.

About

Pluggable memory engine for AI agents — worn first-person doc + event-driven anticipation + sleep-time writer + hybrid retrieval. BYO database / LLM / embedder.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages