Connect evidence, improve what answer engines can understand, and track observed citation share.
Architecture · Backend · Frontend · Invariants · MCP setup · Documentation index · Development
CiteLadder is an evidence-grounded growth intelligence platform for making a brand more likely to be recommended and cited by answer engines—without pretending leading indicators prove causality.
Connect · Analyze · Act · Improve / Verify · Track · AEO · AI Visibility · Evidence-Grounded · Open Source
Most growth tooling answers one question and leaves the evidence behind. CiteLadder answers four connected questions over one governed evidence system:
- What does this business currently say and prove?
- What is missing, weak, contradictory, stale, or hard to discover?
- What are customers demonstrably asking and doing?
- What should the business improve, create, and measure next?
The durable differentiator is not the crawler, the dashboard, or the chat box. It is a project-specific evidence trail that compounds through every crawl, import, generation, audit, and later measurement.
Five user-facing stations sit over four durable capability owners.
| Layer | Owns |
|---|---|
| Site Health | Crawls the owned site, classifies pages structurally, applies page-type-correct checks, and persists comparable-crawl changes |
| Content Intelligence | Turns verified gaps into strategies, briefs, drafts, and post-publication verification |
| Demand Intelligence | Connects GSC, GA4, journeys, prompts, AI Visibility, and later paid marketing evidence to decide what improves next |
| Agent | Answers questions and produces reviewable outputs through the same read tools as the MCP server, with frozen context and reproducible provenance |
AI Visibility is the Track station. The primary measured outcome is increased observed mention/citation share under comparable portfolio and engine conditions; Site Health and demand metrics remain leading indicators.
The Agent is a real layer — the one you spend the most time in — but it owns no data. It is deliberately not a fourth database, an unrestricted chat interface, or an autonomous publisher.
The system runs itself. You are asked exactly twice:
| Decision | Why you are here |
|---|---|
| Ask the Agent and declare what you ship | Agent outputs are the only durable outward-facing deliverables. You choose what to ask for, edit it, and declare what you implemented. |
| Run and schedule audits | Crawls, syncs, and answer-engine audits cost money and hit external systems. You choose when they run. |
Everything else — bounded crawling, structural page classification, deterministic analysis, demand signals, prompt generation, prioritization, and roadmaps — may run automatically within configured limits. Every output records its source and version so stale analysis can be recomputed rather than silently reinterpreted.
immutable evidence
+ versioned page-kind analysis
+ persisted demand and content projections
+ explicit user decisions
Persistence means observed, never automatically true. Site Health keeps immutable acquisition evidence and versioned analysis. Generated content never becomes a fact on its own. The former industry-pack and knowledge-kernel runtime was removed during simplification; page analysis now uses a small generic structural taxonomy with page-type-specific schema and content rules.
owned domain, documents, and integrations
-> immutable evidence
-> page and document understanding
-> project facts, gaps, and demand signals
-> prioritized opportunities
-> brief
-> generated content you edit and save
-> recrawl, resync, or visibility audit
-> before/after observation
-> next recommended action
Every stage is inspectable and versioned, and a later observation never rewrites earlier evidence. Unknown fees, dates, prices, policies, and regulated claims are requested or omitted — never invented. Structured data mirrors saved visible content; it is never a substitute for it.
Stated as plainly as the features, because it is a design constraint rather than a disclaimer:
- no causal conversion diagnosis without adequate behavioural evidence;
- no autonomous publishing or external mutation;
- no model trained on private customer data;
- no automatic sharing of one customer's facts with another;
- no single universal score that hides coverage or industry differences.
unavailable, not configured, and observed zero stay three different things.
From a clean clone, copy the local template and use one Compose command. It builds and starts
Postgres, runs the migration job, then starts the native API, the runner loop and the three
Workers as production serves them: marketing on port 3000, the app on 3001 and docs on 4322.
The browser stays same-origin: the app Worker proxies relative /api/* calls to the API over the
Compose network.
Important: exported
POSTGRES_*andDATABASE_URLshell variables take precedence over Compose's env file. Use theenv -u …command verbatim; it also explicitly selects the copied env file.
# 1. Copy the local-only template. Edit it before enabling integrations or using non-local secrets.
cp .env.example .env
# 2. From the repository root, build and start the complete local stack.
env -u POSTGRES_PASSWORD -u POSTGRES_USER -u POSTGRES_DB -u DATABASE_URL \
POSTGRES_PASSWORD=citeladder_dev_password \
docker compose --env-file .env -f docker-compose.yml \
up -d --build --force-recreate
# 3. Once `docker compose ... ps` shows the services healthy, verify the application and API.
curl -fsS http://127.0.0.1:3000/
curl -fsS http://127.0.0.1:3001/health
curl -fsS http://127.0.0.1:8100/healthNo host-side migration command or separate frontend dev server is needed for this Compose path. Register a user (a workspace is created automatically), create a project, then connect a BYOK provider for Visibility audits or open Site Health to discover and analyze the site.
Commands, environment, entitlement, migration, and the clean-clone runbook:
docs/DEVELOPMENT.md. Release gates are in
docs/release-checklist.md.
| Document | What it covers |
|---|---|
AGENTS.md |
Mandatory implementation rules and the task-specific document map |
docs/architecture.md |
Canonical target product architecture |
docs/invariants.md |
The review-blocking rules |
docs/README.md |
Runtime owners, active work, operations, and historical evidence |
docs/site-health.md |
Site crawl, page kinds, rules, issues, readiness, and crawl changes |
docs/design.md |
Design tokens, screen geometry, and the insight object |
Archived history, when present, is not an implementation authority.
frontend/apps/ marketing, product and docs delivery on Cloudflare Workers
frontend/services/api/ TypeScript API, workers, business logic and operators on Cloud Run
frontend/services/api/src/config/ native application policy
frontend/services/api/migrations/ the single SQL schema baseline (pre-launch)
frontend/packages/contracts/ shared TypeScript API contracts
docs/README.md sole active documentation index
docs/plans/ consolidated backlog and current-work index
docs/archive/ historical plans and evidence; no active dependencies
docs/decisions.md accepted cross-feature decisions and rationale
PostgreSQL owns durable data and queues. Google Cloud hosts the application runtime
and database in the United States (us-central1); the SQL baseline is the only schema author.
Development owns setup, affected-owner completion commands, retry semantics and CI/release diagnostics. Documentation-only work uses textual/reference checks unless a document is packaged or executable input. Do not run broad suites merely to validate prose.
Read AGENTS.md and the owning architecture document before changing code.
CONTRIBUTING.md covers workflow and ownership; Review.md covers
the review checklist and recurring anti-patterns.
CiteLadder is a dirty, active, multi-workstream repository. Preserve unrelated user-owned changes
and follow the repository's task-level validation policy.