Skip to content

About

Graph-based context engine for AI agents

Resources

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

ContextMarmot — a marmot overlooking a glowing underground knowledge graph

ContextMarmot logo ContextMarmot

A graph-based memory engine for agentic systems.

Agents query and write to a persistent knowledge graph via [MCP](https://modelcontextprotocol.io/), getting structured context instead of reading raw files.
Agent  -->  MCP Server  -->  Embed + Search + Traverse + Compact  -->  XML context

Nodes are Obsidian-compatible markdown files with YAML frontmatter and [[wikilinks]]. Open the vault in Obsidian or use the built-in web UI for graph visualization.


Screenshots

Multi-namespace graph with bridge arcs

Multi-namespace graph with bridge arcs

Chat-driven graph curation

Chat-driven graph curation

Node detail panel

Node detail panel with edges, tags, and summary

Folder grouping with contour hulls

Folder grouping with organic contour hulls


Features

  • Graph-based context --- nodes connected by typed, directed edges (structural and behavioral)
  • Structural acyclicity --- DAG-enforced for contains, imports, extends, implements; cycles allowed for behavioral edges
  • Hybrid search --- BM25 full-text (SQLite FTS5) fused with embedding KNN via weighted Reciprocal Rank Fusion, then graph traversal expansion
  • Impact mode --- mode: "impact" (MCP) or marmot query --mode impact treats the query as an exact node ID and returns its ranked inbound blast radius --- the nodes that depend on it, including cross-vault dependents
  • Automatic source watching --- the serve daemon owner watches your indexed source tree and re-indexes changed files (on by default; watch_sources: false opts out)
  • Unresolved-ref resolution --- code references the indexer can't resolve are parked in a refs.db sidecar and retried on later runs; resolutions materialize provenance-stamped edges
  • Token-budget compaction --- results fit your context window, with full/compact/truncated tiers
  • MCP server --- 6 tools over stdio (context_query, context_write, context_tag, context_verify, context_delete, context_namespace)
  • CRUD classifier --- auto-classifies writes as ADD/UPDATE/SUPERSEDE/NOOP using embeddings + optional LLM
  • Multi-namespace --- project isolation with bridge manifests for cross-namespace and cross-vault edges
  • Warrens --- mount selected project graphs from a shared git-backed Warren for multi-repo context
  • Domain tags --- many-to-many semantic categorization; bulk-tag via search query; graph clustering by tag
  • Static analysis indexer --- parses Go (full AST), TypeScript, and 30+ languages into graph nodes with typed edges
  • Graph visualization --- embedded D3 web UI with filters, search, heat overlay, folder grouping, and bridge arcs
  • Graph Curator --- chat-driven curation UI with NL queries, slash commands, and node-ref pills
  • Summary engine --- auto-generates namespace summaries via LLM; regenerates on significant changes
  • Heat map --- co-access frequency tracking with exponential decay; hot edges get traversal priority
  • Integrity verification --- hash-based staleness detection, dangling edge checks, cycle detection
  • Single binary --- Go, zero CGo, zero runtime dependencies

Quick Start

Prerequisites

  • Go 1.21+ (brew install go or go.dev/dl)
  • Node.js 20.19+ (or 22.12+) only if building/running the web UI (marmot ui)
  • (Optional) An OpenAI API key for semantic search. Without one, the mock embedder provides lexical-overlap search.

Build

git clone https://github.com/nurozen/context-marmot.git
cd context-marmot
make build

This produces bin/marmot.

If you also want the embedded web UI:

make build-full

Try the demo

# Seed a demo vault with 8 nodes (auth + database graph)
cd testdata/demo
bash seed.sh

# Index nodes into the embedding store
../../bin/marmot index --dir .marmot

# Query the graph
../../bin/marmot query --dir .marmot --query "user authentication login"

# Verify integrity
../../bin/marmot verify --dir .marmot

# Launch graph UI
../../bin/marmot ui --dir .marmot --no-open

# Open in Obsidian (optional --- open testdata/demo/.marmot as a vault)

Then open http://localhost:3274.

Use in your own project

cd your-project

# Initialize vault, configure embeddings, and set up MCP configs --- all in one step
marmot init

marmot init runs three stages:

  1. Creates the .marmot/ vault directory
  2. configure --- prompts for embedding provider, model, API key, and CRUD classifier (provider + model)
  3. setup --- detects your tools (Claude Code, Codex, VS Code, Cursor) and writes MCP configs

After init, index your source code:

marmot index

Agents automatically connect to the MCP server --- no manual server start needed.

Manual setup (if needed)

Run marmot configure or marmot setup individually at any time:

# Re-configure embedding provider/model/key
marmot configure

# Regenerate MCP configs (e.g., after installing a new tool)
marmot setup

# Target a specific tool
marmot setup --claude
marmot setup --codex
marmot setup --vscode
marmot setup --cursor

Global setup (one config for every project)

marmot setup --global writes user-scope MCP configs instead of project-local ones. The registered command is a bare marmot serve with no vault path embedded — the vault is resolved per-project at serve time via the reverse-route table and .marmot-vault pointer (see Vault discovery), so one global entry covers every registered project on the machine:

marmot setup --global              # auto-detect installed tools
marmot setup --global --claude     # target specific tools
marmot setup --global --dry-run    # print target files + payloads, write nothing

Target files per harness:

Harness User-scope file
Claude Code ~/.claude.json (mcpServers key)
Codex ~/.codex/config.toml ([mcp_servers.context-marmot], idempotent append)
Cursor ~/.cursor/mcp.json (mcpServers key)
VS Code user settings.json (mcp.servers key; macOS ~/Library/Application Support/Code/User/settings.json, Linux ~/.config/Code/User/settings.json, Windows %APPDATA%\Code\User\settings.json)

Existing files are merged non-destructively (unrelated keys and sibling MCP servers are preserved). If a target file exists but cannot be parsed as JSON (including VS Code JSONC settings with comments), setup refuses with an error instead of clobbering it — fix the file or add the server manually.

Codex CLI users: Codex (observed with v0.143) gates MCP servers behind its own approval/trust model. In headless runs (codex exec), project-level .codex/config.toml MCP servers may be hidden, and MCP tool calls are auto-cancelled ("user cancelled MCP tool call") under the default read-only sandbox. Run codex with --dangerously-bypass-approvals-and-sandbox (or mark the project as trusted) for the context-marmot tools to be listed and executable. This is Codex approval behavior, not a marmot limitation.

Mount a Warren for multi-project context

A Warren is a git-backed collection of project .marmot/ vaults. It lets a local workspace query selected upstream project graphs without copying every project graph into the local vault.

# In your local project first: give the workspace a vault_id — the identity
# key for warren bridges involving your own project
marmot configure --vault-id my-project

# From the Warren repository
marmot warren init --id product-platform
marmot warren project import project-a ../project-a/.marmot --vault-id project-a-vault
marmot warren bridge add project-a project-b --relations calls,reads,references
marmot warren doctor

# From your local project or virtual monorepo
marmot warren register product-platform /path/to/product-warren
marmot warren mount --warren product-platform project-a project-b

# Or consume from a git URL via the shared cache (register-free; see docs/warrens.md)
marmot warren add https://github.com/acme/product-warren --id product-platform
marmot warren sync product-platform

# Make one mounted project editable in this workspace
marmot warren edit --warren product-platform project-a

# Cache selected project graphs locally for offline use (burrow always materializes)
marmot warren burrow --warren product-platform project-b

# Undo: deactivate projects, delete caches, remove the Warren
marmot warren unmount --warren product-platform project-a
marmot warren burrow --drop --warren product-platform --all
marmot warren unregister --warren product-platform

Mounted projects are dormant until marmot warren mount activates them (a bare mount/burrow with no project IDs requires an explicit --all). Active Warren projects are included in MCP/CLI graph queries and appear as a separate Warren <id> view in marmot ui. Local namespace views such as default remain local-scoped.

Cross-vault resolution uses a global routing table at ~/.marmot/routes.yml (populated by marmot route add and marmot warren register). Every marmot query/marmot serve loads it, so even a brand-new vault reports vault registry: N remote vaults registered if other vaults are registered on the machine. Set MARMOT_ROUTES=off (or none/0) to run without the global table — useful for hermetic tests and scratch vaults — or point MARMOT_ROUTES=/path/to/routes.yml at an alternate table.

See the "Quickstart: zero to first bridge" walkthrough in docs/warrens.md for the full zero-to-first-bridge flow (including workspace identity), plus Warren layout, read/write policy, authoring commands, bridge policy, materialization, and UI/API behavior.

MCP Tools

Once connected, agents get six tools:

Tool Description
context_query Search the graph by natural language. Returns XML-compacted subgraph within token budget. With mode: "impact", treats the query as an exact node ID and returns its ranked inbound blast radius — the nodes that depend on it.
context_write Write or update a node. Enforces structural acyclicity. Updates embedding index. Accepts optional tags.
context_tag Bulk-tag nodes by semantic search query. Finds nodes matching a query and applies the given tags.
context_verify Check node staleness, dangling edges, and structural integrity.
context_delete Soft-delete (supersede) a node. Excluded from future queries by default.
context_namespace List, create, update, doctor, or remove namespace manifests for the vault.

Serve and the single-owner daemon

Each MCP client spawns its own marmot serve process, so several serves often share one vault. Daemon mode makes that safe: the first serve owns the vault (one engine, one summary scheduler, one graph watcher) and every later serve transparently relays its stdio MCP session to the owner over a unix socket, so all clients see one consistent, always-fresh graph. If the owner dies, a surviving serve re-elects itself and takes over.

Ownership is per-vault, not per-machine: serves pointed at different vaults (e.g. one per project) each elect their own owner and never interact; only serves sharing a vault join the same daemon. Election locks on the vault's daemon.lock inode, so relative/absolute/symlinked spellings of the same vault path all join the same election.

  • On by default: a plain marmot serve joins the per-vault election; no configuration needed.
  • Opt out with marmot serve --no-daemon or MARMOT_NO_DAEMON=1 to run standalone (each serve keeps its own engine over the shared SQLite WAL). Windows always runs standalone.
  • NFS caveat: election uses flock(2) on .marmot/.marmot-data/daemon.lock, whose semantics are unreliable on NFS/network filesystems — set MARMOT_NO_DAEMON=1 for vaults that live on one.

CLI Reference

Command Description
marmot init [--dir .marmot] Create a new vault, run configure, then setup
marmot configure [--dir .marmot] Interactive prompt for embedding provider, model, API key, and CRUD classifier
marmot setup [--dir .marmot] [--global] [--dry-run] [--claude] [--codex] [--vscode] [--cursor] Generate MCP configs for detected (or specified) tools. --global writes user-scope configs with bare marmot serve (no vault path); --dry-run previews without writing.
marmot index [--dir .marmot] [--force] [<path>] [--incremental] Index node files or run static analysis on source code. --force rebuilds all embeddings.
marmot query --query "..." [--dir .marmot] [--depth 2] [--budget 12288] [--mode adjacency|impact] Query the knowledge graph. --mode impact treats the query as an exact node ID and returns its ranked inbound blast radius
marmot verify [--dir .marmot] Run integrity and staleness checks
marmot serve [--dir .marmot] [--no-daemon] Start the MCP server on stdio (see serve and the single-owner daemon)
marmot status [--dir .marmot] Show vault stats: node counts, edges, embeddings, namespaces, heat map
marmot watch [--dir .marmot] Standalone source + vault watcher for auto-reindex. Refused while a serve daemon owner is alive (the owner already watches sources when watch_sources is on)
marmot namespace create/list/update/doctor/remove ... Manage per-namespace _namespace.md manifests
marmot bridge <ns-a> <ns-b> [--relations ...] Create bridge manifest between two namespaces
marmot warren init --id <id> [--warren-dir .] Create a Warren repository manifest
marmot warren project import <project-id> <source-.marmot> [--vault-id <id>] Copy an existing project vault into a Warren
marmot warren project add <project-id> --path <project-.marmot> [--vault-id <id>] Register an already placed project vault in a Warren
marmot warren project list/remove/rename ... List or maintain Warren project entries
marmot warren project set-readonly <project-id> [--off] Author-side write policy: veto consumer edits for one project
marmot warren bridge add/list/remove ... Maintain Warren-owned project bridge policy
marmot warren doctor [--workspace] [--json] Validate a Warren repository (or, with --workspace, this workspace's warren state)
marmot warren format Normalize a Warren manifest
marmot warren register <id> <path> Register a git-backed Warren repository in this workspace
marmot warren list [--json] List registered Warrens and local mount state
marmot warren mount --warren <id> <project-id>...|--all Activate selected (or, with --all, every) Warren project for query/UI use
marmot warren unmount --warren <id> <project-id>...|--all Deactivate Warren projects (non-destructive; burrow caches kept)
marmot warren burrow --warren <id> <project-id>...|--all Activate and cache Warren projects under .marmot-data/ (always materializes)
marmot warren burrow --drop --warren <id> <project-id>...|--all Delete burrow caches (clears the materialized flag with the last one)
marmot warren unregister --warren <id> [--force] Remove a Warren from the workspace (refuses while mounts/caches exist unless forced)
marmot warren status --warren <id> [--json] Show registered/active/editable/materialized project state (flags unreachable checkouts)
marmot warren edit [--off] --warren <id> <project-id> Toggle write access for one Warren project (edit implies mount)
marmot warren refresh [--pull] --warren <id> Reload Warren state for live engines; --pull also fast-forwards the checkout and re-materializes stale burrow caches
marmot warren propose [--warren <id>] [<project-id>] Branch + commit one project's editable-mount edits for review (local-only; never pushes)
marmot summarize [--namespace ...] Force summary regeneration for a namespace
marmot reembed [--dir .marmot] Regenerate all embeddings (use after changing provider/model)
marmot sdk [--out ./marmot-sdk.ts] Generate a type-safe TypeScript SDK from MCP tool schemas
marmot ui [--dir .marmot] [--host 127.0.0.1] [--port 3274] [--no-open] Start the embedded graph visualization UI (binds loopback only by default; --host 0.0.0.0 exposes it)

Configuration

_config.md frontmatter knobs (all optional; marmot configure manages the embedding/classifier ones):

Key Default Description
token_budget 12288 Default token budget for query responses
search_bm25_weight 1.0 RRF weight of the lexical BM25 arm. search_bm25_weight: 0 is the kill switch that drops the lexical arms (see the note below --- it does not restore pre-hybrid ranking)
search_vector_weight 1.0 RRF weight of the vector KNN arms; 0 disables them
search_autocut false Opt-in score-gap autocut on the fused result list
search_autocut_gap 0.2 Autocut jump ratio in (0, 1]
search_max_results 20 Entry-candidate pool size --- the final entry cap; each retrieval arm over-fetches beyond it before fusion
search_min_score (off) Similarity floor applied per vector arm before fusion, in the vector index's own score space --- 1/(1 + L2 distance), not cosine (see the note below); non-positive disables it
watch_sources true Automatic source watching by the serve daemon owner; explicit false opts out
source_root (unset) Last source root indexed by marmot index <src> --- recorded automatically and used as the source watcher's root (v1 watches exactly one root)

What search_bm25_weight: 0 does and does not do. It disables the lexical (BM25) arms --- local and per-remote-vault --- so only the vector KNN arms feed retrieval. It is not a rollback to pre-hybrid behavior: entry candidates are still over-fetched per arm and cut to search_max_results after fusion, the surviving vector arms are still combined by reciprocal-rank fusion (one arm per vault, so cross-vault results are merged by rank rather than by raw --- and incomparable --- similarity scores), the reported score is an RRF score rather than a raw vector similarity, and search_min_score / search_autocut still apply when set. Use it when the lexical arm is hurting your ranking; use a downgrade of the binary if you need the exact pre-hybrid result ordering.

What scale search_min_score is on. The vector index scores a hit as 1 / (1 + d) where d is the L2 distance between the query and the stored embedding --- reciprocal distance, not cosine similarity. For unit-length embeddings d = sqrt(2 - 2c) for cosine c, so the score of cosine c is 1 / (1 + sqrt(2 - 2c)): the scale runs from 1.0 (identical) down to about 0.333 (opposite vectors), and the midpoint 0.5 is exactly cosine 0.5. Set the floor in that space --- 0.5 is a moderate cut, 0.6 (cosine ~0.69) is aggressive, and 0.8 means cosine ~0.97, which will drop nearly everything. Values above 1.0 disable the vector arms entirely.

Architecture

cmd/marmot/              CLI (init, configure, setup, index, query, serve, verify)
internal/
  config/                Vault config, .env key storage, embedder + classifier factory
  node/                  Markdown parser/writer, atomic file I/O, temporal fields
  graph/                 In-memory graph, adjacency lists, cycle detection
  verify/                Hash integrity, staleness, structural checks
  embedding/             SQLite store, KNN search, FTS5 BM25 index, OpenAI + mock embedders
  traversal/             BFS traversal, token-budget XML compaction
  llm/                   LLM provider interface (OpenAI, Anthropic, mock)
  classifier/            CRUD classifier: embedding + LLM + distance fallback
  namespace/             Namespace manager, bridge manifests, qualified ID resolution
  warren/                Git-backed multi-project Warren manifests and local mount state
  summary/               Namespace summary generation, async scheduler
  update/                Source change detection, staleness propagation, source watcher
  indexer/               Static analysis: Go AST, TypeScript regex, generic, runner
  refs/                  Unresolved-reference parking sidecar (refs.db), retry + resurrection
  mcp/                   MCP server (6 tools), engine wiring

See docs/architecture.md for the full system design.

Current Status

MVP complete. Evaluated on SWE-QA --- 20 code comprehension questions across django, flask, pytest, and requests:

Metric Vanilla (file tools) Hybrid (ContextMarmot) Improvement
Answer quality (1-5) 4.62 4.62 identical
Tokens per question 151,327 95,876 -37%
Cost per question $0.1065 $0.0834 -22%
Avg turns 7.5 6.9 -8%

Same quality. Lower cost. The graph acts as a navigation map --- agents query it first, then read only the files and line ranges it identifies, skipping broad exploration.

See docs/benchmark.md for the full per-question breakdown and methodology.

Known MVP limitations

  • Go-side KNN --- search scans all embeddings in Go. Fine for thousands of nodes; would need sqlite-vec for 100k+.
  • Limited provider selection --- OpenAI and mock are supported. Voyage AI, Ollama, and other providers are not yet available.

Post-MVP roadmap

See docs/implementation_plan.md for the full plan and phase status.

Documentation

Document Description
Architecture Full system design and component interactions
Bridges Namespace and cross-vault bridge configuration
Warrens Git-backed multi-project graph mounting and edit policy
Dens Central dens under $MARMOT_HOME: create/adopt, routes, discovery, link modes (edit/pinned/live), contribute, bridges, schema:1 JSON
Embedding Providers Embedding provider setup and fallback behavior
CRUD Classifier Write classification (ADD/UPDATE/SUPERSEDE/NOOP)
TypeScript SDK Type-safe SDK generation and usage
Development Build commands, node format, and edge types
Benchmark SWE-QA evaluation methodology and results
Data Structures Node, edge, and vault format specifications
Implementation Plan Full roadmap with phase status

License

Apache 2.0 --- see LICENSE.

About

Graph-based context engine for AI agents

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages