Research into a universal, language-agnostic representation that lets next-generation agents understand and operate on codebases without retrieving the underlying source.
Still figuring this out. We're just building, experimenting, chasing weird ideas, and seeing where they take us. Thanks for sticking around.
This project is still very much alive and under active development. Some parts work. Some parts don't. Some parts work until they suddenly don't.
Expect:
🧪 Experimental ideas
🛠️ Things being rebuilt from scratch
💥 Occasional breakage
🌀 Features that may change direction
✨ Unexpectedly good ideas
🤷 The occasional "well, that wasn't supposed to happen"
The goal right now isn't to make everything look finished.
It's to build, experiment, break things, learn, and keep making it better.
If you're looking for something polished and predictable... this might be a little early. 😄
What is the minimum information an agent needs to understand a codebase and acomplish a task without reading unnecesary/unrelated code?
That is the question behind Prime.
Prime is a research project exploring how an entire software repository can be transformed into a single, extremely compact knowledge artifact containing the smallest useful units of information required for an agent to understand, navigate, analyze, and reason about the repository.
Prime does not aim to compress source code.
Prime aims to make the source code more understandable and navigatable for as many agent questions as possible.
CODEBASE
│
▼
┌─────────────────┐
│ PRIME │
│ │
│ derived │
│ knowledge │
│ │
│ minimal units │
│ compact │
│ indexed │
│ language │
│ agnostic │
└────────┬────────┘
│
▼
AGENT
│
┌────────┴────────┐
▼ ▼
ANSWER ACT
A codebase contains enormous amounts of information.
An agent rarely needs all of it.
For a question such as:
Who calls
AuthService.login?
the agent should not need to:
read files
→ search symbols
→ follow imports
→ parse declarations
→ inspect references
→ reconstruct relationships
Prime should already contain the derived knowledge necessary to answer:
AuthService.login
← CheckoutController
← AdminController
← SessionRefreshJob
For:
What does
AuthService.logindepend on?
Prime should already know:
AuthService.login
→ UserRepository.findByEmail
→ PasswordVerifier.verify
→ SessionStore.create
For:
What can
AuthService.loginreturn or throw?
Prime should expose the derived contract:
returns:
Session
may throw:
UserNotFound
InvalidCredentials
The source implementation is not required for these answers.
That is the purpose of Prime.
This is a fundamental constraint.
Prime should never become a compressed copy of the repository.
The final representation is not intended to contain:
- source files
- source snippets
- complete ASTs
- reconstructed implementations
- raw file contents
- duplicated syntax
- arbitrary text copied from the repository
Instead, Prime contains derived knowledge.
Conceptually:
SOURCE
│
│ analyze
▼
DERIVED KNOWLEDGE
│
│ minimize
▼
PRIME
Prime is intentionally lossy.
It should not be possible, or even desirable, to reconstruct the original codebase from Prime.
The source repository remains the authority.
Prime is the distilled knowledge layer above it.
Prime is not a database project.
It is not a graph database project.
It is not a compression project.
It is not a search engine project.
Those may become implementation techniques.
The actual objective is:
Maximize the number and quality of agent questions that can be answered without retrieving source code, while minimizing the amount of information, I/O, computation, latency, and context required.
A useful abstraction is:
AGENT KNOWLEDGE
▲
│
│
┌───────────┴───────────┐
│ PRIME │
│ │
│ minimum useful │
│ representation │
└───────────┬───────────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
less I/O less parsing less searching
│ │ │
└─────────────┼─────────────┘
▼
less agent work
The important metrics are therefore not simply:
file size
node count
query latency
We care about:
questions answerable without source
useful knowledge / byte
useful knowledge / token
useful knowledge / I/O
retrieval latency
computation required
context required
agent tool calls avoided
Prime ultimately asks:
What is the smallest useful representation of a software repository that preserves enough information for a next-generation agent to understand, navigate, modify, and reason about that repository without retrieving the underlying code?
That question contains several smaller questions:
1. What questions do agents ask?
2. What information is required to answer those questions?
3. What information can be derived from the codebase?
4. Which derived information is actually useful?
5. What is the smallest useful knowledge unit?
6. Which information is redundant from an agent's perspective?
7. How can those units be represented compactly?
8. How can they be retrieved with minimal I/O?
9. How can the representation remain language agnostic?
10. How can it scale from tiny repositories to enormous monorepos?
11. How can it remain useful across different agent architectures?
12. What happens when the available analysis is incomplete or uncertain?
Prime is designed specifically around the architecture of modern and emerging coding agents.
The agent is not a passive reader.
A modern agent typically operates in a loop:
observe
↓
reason
↓
retrieve
↓
reason
↓
act
↓
observe
↓
...
Prime therefore needs to optimize for the information flow through that loop.
We need to understand:
- model context
- attention
- context windows
- context caching
- tool use
- agent memory
- retrieval loops
- planning
- tool schemas
- structured tool results
- progressive disclosure
- external memory
- agent failure modes
- information overload
- repeated retrieval
The goal is not merely to give an agent more information.
The goal is to give it the right information in the smallest useful form at the right time.
Prime should investigate the smallest independently useful unit of codebase knowledge.
It may turn out to be:
an entity
a fact
a relationship
a contract
a behavioral fact
a dependency
a state transition
a provenance record
or some new primitive we have not identified yet.
A conceptual example:
AuthService.login
── CALLS ──>
UserRepository.findByEmail
Another:
AuthService.login
── RETURNS ──>
Session
Another:
AuthService.login
── MAY_THROW ──>
InvalidCredentials
The research must determine whether units like these are sufficient, how they should be combined, and how they can be encoded with minimal overhead.
The concept is intentionally called a knowledge unit, not a graph node, because Prime must not assume its final physical representation in advance.
Prime is fundamentally a semantic distillation problem.
The transformation is:
SOURCE
│
┌───────────┼───────────┐
▼ ▼ ▼
syntax semantics structure
│ │ │
└───────────┼───────────┘
▼
DERIVED FACTS
│
▼
REMOVE REDUNDANCY
│
▼
MINIMUM USEFUL FORM
│
▼
PRIME
The important distinction is:
compression:
make the same information smaller
Prime:
remove information that is unnecessary
while preserving what the agent needs
Prime therefore requires research into information theory, semantic compression, sufficient statistics, information bottlenecks, rate-distortion, and other ways of understanding what information is actually necessary.
Prime must work across programming languages.
This is a hard requirement.
It should be able to process:
TypeScript
JavaScript
Python
Rust
Go
Java
Kotlin
C
C++
C#
Ruby
PHP
Swift
Scala
Dart
Lua
and others
It must also handle polyglot repositories:
frontend/
TypeScript
backend/
Rust
services/
Go
native/
C++
automation/
Python
infrastructure/
Terraform
YAML
The representation must preserve meaningful relationships across those boundaries.
For example:
TypeScript
│
▼
HTTP API
│
▼
Rust service
│
▼
gRPC
│
▼
Go service
│
▼
database
The universal model should not attempt to make every language identical.
Instead:
language-specific analysis
│
▼
universal semantics
│
▼
PRIME
Language-specific frontends may understand different concepts with different levels of precision.
Prime should make those differences explicit.
Not every language exposes the same amount of static information.
Dynamic dispatch, reflection, macros, metaprogramming, generated code, runtime loading, and other features can make some facts impossible to establish statically.
Prime therefore needs to distinguish knowledge such as:
EXACT
DERIVED
INFERRED
UNKNOWN
and, where appropriate, preserve:
confidence
provenance
analysis source
source location
revision
Example:
AuthService.login
CALLS
UserRepository.findByEmail
confidence:
exact
evidence:
static symbol resolution
versus:
PluginManager
MAY_CALL
PaymentProvider
confidence:
inferred
reason:
dynamic dispatch
Prime must never pretend uncertain knowledge is exact.
Prime must work across radically different repository sizes.
5 files
↓
500 files
↓
50,000 files
↓
500,000 files
↓
millions of files
A tiny project should not require a heavyweight infrastructure stack.
A huge repository should not require loading the entire knowledge representation into memory.
The design must investigate:
- streaming analysis
- incremental processing
- parallel analysis
- partial loading
- memory mapping
- immutable structures
- compact indexing
- sharding where useful
- content addressing
- caching
- incremental invalidation
Scale is not an optimization phase to add later.
It is part of the core problem.
Prime should build on existing research rather than pretending the field starts here.
Important systems include:
A language-agnostic source-code indexing protocol covering symbols, definitions, references, and implementations.
https://github.com/sourcegraph/scip
A persistent representation of language-server information designed to make code intelligence available without repeatedly running the language server.
https://github.com/microsoft/lsif-node
A rich program representation combining multiple forms of program structure and analysis.
https://github.com/joernio/joern
An incremental, error-tolerant parsing ecosystem supporting many languages.
https://github.com/tree-sitter/tree-sitter
A practical demonstration that a compact structural map of a repository can provide substantial value to an LLM while remaining within a token budget.
https://aider.chat/2023/10/22/repomap.html
Projects such as codebase-index and other repository intelligence systems explore hybrid combinations of symbol indexes, lexical search, relationships, embeddings, and context assembly.
Prime should study all of these carefully, including their source code and architectural tradeoffs.
Prime deliberately has a very wide research scope.
The optimal solution may come from outside traditional code indexing.
We will investigate:
- entropy
- information bottleneck
- sufficient statistics
- minimum description length
- rate-distortion
- semantic compression
- information-preserving transformations
- delta encoding
- variable-length integers
- compressed adjacency
- grammar compression
- succinct graphs
- WebGraph-style techniques
- Elias-Fano
- rank/select
- bit packing
- inverted indexes
- FSTs
- minimal perfect hashing
- learned indexes
- approximate indexes
- Bloom filters
- quotient filters
- Roaring bitmaps
- content addressing
- Merkle trees
- Merkle DAGs
- peer-to-peer distribution
- deduplication
- distributed caching
- synchronization
- CRDTs
- integrity
- provenance
- signed knowledge
- commitments
- Merkle proofs
- searchable encryption
- privacy-preserving retrieval
- zero-knowledge techniques
Cryptography is not being researched because encryption is inherently faster.
It is being researched because trust, provenance, identity, deduplication, and distributed knowledge verification may materially improve the overall system.
- memory mapping
- page cache behavior
- SSD / NVMe I/O
- CPU cache locality
- SIMD
- zero-copy access
- concurrency
- immutable data structures
- persistent data structures
- context management
- memory
- tool use
- model attention
- context caching
- retrieval loops
- agent planning
- context selection
- long-context behavior
- structured tool outputs
Prime intentionally does not restrict its research to technologies traditionally associated with code intelligence.
The eventual Prime product is expected to produce a single logical knowledge artifact representing a codebase.
The physical representation is deliberately undecided during research.
It could eventually use:
- one binary file
- a custom format
- memory-mapped structures
- compressed arrays
- content-addressed blocks
- specialized indexes
- another structure not yet identified
The constraint is the outcome, not the implementation:
The artifact should expose the maximum useful codebase knowledge with the minimum retrieval cost.
The final representation must never require the original source to answer questions that can already be derived from Prime.
Prime should eventually support questions such as:
Where is X?
What is X?
What does X represent?
What does X depend on?
What depends on X?
Who calls X?
What does X call?
What implements X?
What does X implement?
What references X?
What uses X?
What tests X?
What configuration affects X?
What is the architecture around X?
What components are connected to X?
What would be affected if X changes?
What are the contracts of X?
What are the known behaviors of X?
What are the known failure modes of X?
What is the smallest context required to understand X?
These are examples, not yet the final API.
The research must discover the actual question space of next-generation coding agents.
Traditional systems often begin with:
nodes
edges
rows
documents
chunks
vectors
Prime should begin with:
agent question
↓
information requirement
↓
minimum knowledge
↓
retrieval
The underlying representation is secondary.
The end goal is an efficient answer.
The defining property of Prime is:
Agent question
↓
Prime
↓
answer
not:
Agent question
↓
Prime
↓
source lookup
↓
parse
↓
search
↓
answer
Prime may retain provenance pointing back to the source, but provenance is not the answer.
The source is the authority and fallback.
Prime is the precomputed knowledge.
Prime should be intentionally non-reversible.
The transformation:
source → prime
should discard information that does not contribute useful knowledge for agents.
This can include:
- formatting
- repeated syntax
- implementation details that are irrelevant to known questions
- redundant representations
- incidental identifiers
- non-semantic text
- source-level structure that does not improve retrieval
This is a design advantage, not a limitation.
Prime follows one rule:
Do not design the answer before understanding the problem.
Research should use:
- primary documentation
- official specifications
- source repositories
- academic papers
- implementation analysis
- benchmarks
- experiments
- real repositories
- agent evaluations
Every major conclusion should distinguish between:
FACT
OBSERVATION
EXPERIMENTAL RESULT
HYPOTHESIS
INFERENCE
OPEN QUESTION
Do not present an architectural preference as an established fact.
Do not choose a technology because it is familiar.
Do not preserve a design simply because we implemented it first.
The repository itself is deliberately organized into three layers.
RESEARCH
│
▼
RESEARCH FINDINGS (specs)
│
▼
DOCS
│
▼
PRIME
Contains the external knowledge and research.
It answers:
What do we know?
Contains the technical conclusions derived from the research.
It answers:
What should Prime be?
Contains the project context, standards, agent roles, and constraints required to work on Prime effectively.
It answers:
How should agents work on Prime?
Prime/
│
├── README.md
│
├── specs/
│ ├── agents/
│ ├── codebase-analysis/
│ ├── prior-art/
│ ├── representation/
│ ├── information-theory/
│ ├── compression/
│ ├── indexing/
│ ├── retrieval/
│ ├── storage/
│ ├── distributed/
│ ├── cryptography/
│ ├── systems/
│ ├── languages/
│ ├── experiments/
│ ├── benchmarks/
│ └── references/
│
├── docs/
│ ├── research-synthesis/
│ ├── requirements/
│ ├── representation/
│ ├── retrieval/
│ ├── architecture/
│ └── decisions/
│
└── .acc/
├── config.yaml
├── config/
└── agents/
Nothing in this structure is considered final until the research justifies it.
The research should progress in this order.
PHASE 1
Understand next-generation agents
│
▼
PHASE 2
Understand codebase information
│
▼
PHASE 3
Study existing representations
│
▼
PHASE 4
Discover the minimum useful knowledge unit
│
▼
PHASE 5
Research representation and compression
│
▼
PHASE 6
Research retrieval and agent interaction
│
▼
PHASE 7
Research scale and language agnosticism
│
▼
PHASE 8
Build isolated experiments
│
▼
PHASE 9
Benchmark agent tasks
│
▼
PHASE 10
Synthesize findings
│
▼
PHASE 11
Write the Prime specification
│
▼
PHASE 12
Configure ACC
│
▼
PHASE 13
Build Prime
The order matters.
Implementation is the final phase, not the first.
Prime should eventually be evaluated against:
- small TypeScript project
- small Python project
- small Rust project
- small Go project
- large TypeScript monorepo
- large Java monorepo
- large Rust workspace
- large C/C++ codebase
- large Python ecosystem repository
Repositories containing multiple languages with relationships across boundaries.
Repositories containing:
- generated code
- macros
- reflection
- dynamic dispatch
- metaprogramming
- build-generated sources
- unusual module systems
- large dependency graphs
The goal is not merely:
"support many parsers."
The goal is:
derive a useful universal knowledge representation across fundamentally different programming models.
Prime must be designed to investigate repositories ranging from:
tiny
small
medium
large
very large
monorepo
polyglot monorepo
The representation should not assume:
all code fits in memory
all code can be analyzed in one pass
all relationships are local
all languages behave statically
Potential techniques include:
- streaming
- parallel analysis
- incremental analysis
- partial loading
- memory mapping
- compact indexes
- content addressing
- immutable snapshots
- distributed analysis
- caching
The appropriate combination must be established through research and measurement.
The strongest Prime benchmark should ultimately be based on agent tasks, not storage benchmarks alone.
Compare, on identical repositories and tasks:
raw filesystem access
repository maps
existing code indexes
graph retrieval
hybrid retrieval
Prime
Measure:
task success
time
tool calls
bytes transferred
tokens exposed
retrieval precision
retrieval recall
context redundancy
agent corrections
source accesses required
The central question is:
How much useful codebase understanding can an agent obtain before it needs to inspect source?
A successful Prime artifact should dramatically reduce:
source reads
searches
parsing
relationship discovery
context reconstruction
agent tool calls
while preserving:
correctness
useful context
architectural understanding
relationship accuracy
agent task performance
The ideal system is not merely smaller.
It makes the codebase legible.
Prime is designed to work alongside Agent Code Context (ACC).
ACC and Prime solve different problems.
ACC provides project-level agent context:
standards
architecture
rules
contracts
workflows
project knowledge
Prime investigates codebase-derived knowledge:
symbols
relationships
dependencies
architecture
contracts
behavior
impact
Conceptually:
CODEBASE
│
┌────────┴────────┐
▼ ▼
ACC PRIME
│ │
project knowledge derived knowledge
│ │
└────────┬────────┘
▼
AGENT
Prime should eventually become a lower-level knowledge layer that ACC can consume.
ACC should not dictate Prime's internal representation.
Prime should not replace ACC's project-context model.
Prime is not:
- a source-code compressor
- a code archive
- a compiler
- a programming language
- a language server
- a graph database
- a vector database
- an embedding store
- a documentation generator
- an IDE
- an agent
- a replacement for Git
- a replacement for source code
Prime is also not required to use:
- graphs
- databases
- embeddings
- cryptography
- P2P
- memory mapping
- custom binary formats
Those are research areas.
The final system should use whatever combination of techniques best satisfies the objective.
Research / Early Stage
At this stage:
- no final Prime format exists
- no final graph model exists
- no final storage design exists
- no final retrieval API exists
- no final compression algorithm exists
- no final language model exists
Those decisions are intentionally postponed.
The current goal is to understand the problem deeply enough to make them responsibly.
Prime is open research.
The value of the project is not only the eventual implementation.
The research itself should remain useful to other people working on:
- code intelligence
- static analysis
- code search
- AI agents
- programming languages
- databases
- storage engines
- compression
- distributed systems
- information retrieval
- machine-readable software representations
Useful contributions include:
- papers
- technical references
- repositories
- implementation analysis
- benchmarks
- experiments
- counterexamples
- alternative approaches
- failed approaches
- new ideas
- corrections
A finding that disproves a promising idea is useful.
A benchmark showing that a "clever" approach is actually slower is useful.
A technology from an unrelated field that turns out to solve part of this problem is especially useful.
Everything in this repository should ultimately be measured against one sentence:
Prime exists to minimize the information an agent must retrieve from a codebase while maximizing the agent's ability to understand and reason about that codebase.
Not smaller files for their own sake.
Not faster graphs for their own sake.
Not better databases for their own sake.
Not more search results.
Less information retrieved. More understanding achieved.
That is Prime.
Latest benchmark:
| Metric | Result |
|---|---|
| Derivation | 334 ms |
| Artifact size | 1.2 MB |
| Artifact/Source ratio | 1.196 |
| Retrieval p50 (warm) | 169 µs |
| Retrieval p95 (warm) | 305 µs |
| Accuracy | 8.2% |
| Source-free accuracy | 8.2% |
| Entity precision | 0.35 |
| Entity recall | 0.40 |
| Entity F1 | 0.37 |
| Relationship precision | 0.00 |
| Relationship recall | 0.00 |
| Relationship F1 | 0.00 |
| MRR | 0.40 |
| Recall@1 | 0.1% |
| Recall@3 | 0.1% |
| Recall@5 | 0.1% |
| Recall@10 | 0.1% |
Repository: bat (rust, small)
Repository: httpx (python, small)
Repository: express (javascript, small)
Repository: gin (go, small)
Repository: spdlog (cpp, small)
Integrity: ✅ Valid Repos: 5/5 completed Warnings: bat: source savings not measured (requires controlled baseline), httpx: source savings not measured (requires controlled baseline), express: source savings not measured (requires controlled baseline), gin: source savings not measured (requires controlled baseline), spdlog: source savings not measured (requires controlled baseline)
Commit: 0154620eaeb9
Benchmark version: 1.0.0
Environment: macos / aarch64 / Apple M2
Full machine-readable result: benchmarks/results/latest.json
We also tested 5 popular JavaScript/TypeScript frameworks using the disk-perf-git-and-pnpm methodology to understand real-world installation performance.
See the full comparison: benchmarks/results/comparison.md
Quick Summary (3-run averages):
| Framework | Packages | Avg Total (s) | Cold Install (s) |
|---|---|---|---|
| Vite + Vue + TS | 48 | 7.98 | 7.88 |
| SvelteKit | 56 | 8.21 | 7.49 |
| Nuxt.js | 606 | 10.66 | 8.79 |
| Next.js | 360 | 10.85 | 9.23 |
| Remix | 764 | 11.12 | 9.22 |
Lightweight frameworks (Vite+Vue, SvelteKit) install ~3x faster than full frameworks (Remix, Next.js) primarily due to fewer packages.
Full machine-readable result: benchmarks/results/latest.json