Three-agent company-intelligence system. Bounded loops, deterministic grounding, cassette evals, provenance-attested releases.
v1.1.0. Deployed and live on a single Hetzner VPS behind a Cloudflare tunnel. Zero open inbound ports.
- Live demo:
https://recon.jakemorganlabs.dev(HMAC-signedPOST /webhook/recon; publicGET /health; rate limit 5 req / 10 s / IP) - Sample dossier:
docs/evidence/sample_dossier.md(a real live run, source-cited) - Eval report:
docs/evidence/eval_report.md(30-case cassette suite, reproducible offline)
The runtime is a deterministic TypeScript orchestrator that systemd runs. It is not the n8n workflows in workflows/. See the architecture note below. No evidence file contains a real secret.
Recon turns a company name into a structured, source-cited intelligence dossier.
- One bounded extraction captures the brief.
- A Search agent gathers evidence from the open web.
- An Analyst agent extracts signals from that evidence.
- A Synthesis agent writes the dossier.
- The agents coordinate through inspectable shared state.
Accountable orchestration is the defining property. Each agent gets a minimal, role-specific toolset. A schema gate checks each output before the next agent reads it. Every claim traces to a retrieved evidence snippet. The system takes no autonomous action on the world. The worst case is a dossier with explicit gaps, never a fabricated claim and never an action taken.
flowchart LR
subgraph Brief["Brief Normalizer"]
A[Raw Request] --> B[Structured Brief]
end
subgraph Search["Search Agent (bounded web)"]
B --> C[Web Queries]
C --> D[Page Fetches]
D --> E[Evidence Store]
end
subgraph Analyst["Analyst Agent (no web)"]
E --> F[Signal Extraction]
F --> G{Coverage Check}
G -->|Unfilled slots & budget| C
G -->|Done| H[Signals]
end
subgraph Synthesis["Synthesis Agent (no web)"]
H --> I[Dossier Draft]
I --> J[Grounding Gate]
end
subgraph Output["Output"]
J -->|Pass| K[Grounded Dossier]
J -->|Fail| L[Gapped Dossier]
end
The no-action posture is the safety property. The system only produces a dossier a human reads. It sends nothing, writes nothing, transacts nothing. The worst outcome of any agent error is a report that flags what it could not verify.
The 1.0 SRS/TDD specified the three agents as n8n AI Agent nodes, with the n8n workflow graph as the orchestrator. The shipped and deployed system is a deterministic TypeScript orchestrator (src/pipeline.ts). The agents, the bounded coverage loop, the grounding gate, and the schema checks are all TypeScript, run as a systemd service behind an HMAC endpoint. This was a deliberate change: cohesion with the other portfolio pieces, and deterministic, testable control flow. The workflows/ directory retains n8n JSON exports as reference triggers only. They are not the running system, and Recon does not need them. The TypeScript core is the source of truth. Rev 1.1 of the SRS/TDD records this decision and its reasons.
A full run outlives the tunnel's edge timeout. The API is therefore asynchronous.
- Send a signed
POST /webhook/recon. The service verifies the signature, creates the run, and returns 202 with arun_id. - Poll
GET /run/<id>until the status is terminal. - Read the dossier from the terminal response.
NOTE: An identical request body derives the same idempotency key and returns the existing run. Vary the target to force a fresh run.
Recon ships with a labeled eval set of 30 cases: 10 rich, 8 thin, 5 empty, 7 adversarial. The suite runs against recorded cassette fixtures. No live web is required. The suite gates CI and releases.
| Suite | Cases | Metric | Threshold | Gate |
|---|---|---|---|---|
| Evidence Recall (recall@k) | 10 rich | >= 0.80 | 1.00 | Hard |
| Structural Validity | 30 all | >= 0.95 | 1.00 | Hard |
| Grounding Integrity (recast-gap rate) | 30 all | <= 0.05 | 0.0000 | Hard |
| Gap Correctness (FAR / FAR_inv) | 10 thin | <= 0.05 / <= 0.20 | 0.00 / 0.01 | Hard |
| Injection Resistance (instructions obeyed) | 7 adversarial | 0 | 0 | Hard |
Result: 30 of 30 cases pass. All gates are green. These are deterministic-replay numbers. The suite runs offline against recorded web cassettes with a deterministic LLM stub, so it is a reproducible regression gate, not a measure of live model quality. Recall and validity at 1.00 partly reflect internally consistent fixtures. The meaningful hardening result is injection resistance: 0 of 7 adversarial instructions obeyed. That metric exercises real control flow. The agents route injected page text as data, never as instructions. Live-quality evidence is the sample dossier in docs/evidence/, produced by a real signed run over the tunnel. The thresholds shown are the gates that evals/run.ts enforces. Regenerate with npm run eval (needs Postgres, no web, DEEPINFRA_API_KEY unset).
Recon's security is structural, not just procedural.
- No-action posture. The worst case is a dossier with explicit gaps. The system never sends an email, writes an external row, or calls an API outside the evidence store.
- HMAC-signed webhooks. Every inbound request is signed with a shared secret. A stale request (over 5 minutes) or an invalid signature is rejected.
- Secrets only in an on-box env file.
deploy/.env.productionis chmod 600, gitignored, and loaded by the systemd service. It never enters git and never leaves the box.scripts/secret_gate.shgreps the tree for token patterns and fails CI on any literal-key hit. - OIDC-attested releases. Every release artifact carries a GitHub-signed provenance attestation via
actions/attest-build-provenance.
Dashboard: a Metabase compose file ships in deploy/metabase/docker-compose.yml but is not deployed. The single 1.9 GB VPS already runs three Node services plus host Postgres. An always-on Java dashboard is the largest OOM risk on that box, so it stays an on-demand operator tool, not a live service. Run-level observability lives in Postgres: the audit, tool_call, and run tables log every stage transition, evidence fetch, and status change per run_id.
make cost-month| Metric | Value | Source |
|---|---|---|
| Cost per run (p50) | not yet metered | scripts/cost_monthly.sh |
| Cost per run (p95) | not yet metered | scripts/cost_monthly.sh |
| Cache savings rate | 0 (no prompt caching on DeepInfra) | scripts/cost_monthly.sh |
Gemma 4 on DeepInfra does not support prompt caching, so the cache columns stay zero. The metric structure is in place for future model swaps.
git clone https://github.com/jakemorganlabs/recon_multiagent.git
cd recon_multiagent
cp .env.example .env
# set your DeepInfra key and search API key for live runs
npm install
docker run -d -e POSTGRES_USER=recon -e POSTGRES_PASSWORD=recon -e POSTGRES_DB=recon -p 5432:5432 postgres:18
npm run migrate
npm test
npm run eval # offline cassette suite; leave DEEPINFRA_API_KEY unset for deterministic replayWARNING: Do not run the eval with a live DEEPINFRA_API_KEY set. A live model rewords the recorded queries, every cassette lookup misses, and the numbers are meaningless. The deterministic mode activates only when the key is unset and CASSETTE_MODE=play.
See docs/runbook.md for redeploy, secret rotation, cassette refresh, DLQ checks, and the closeout protocol.
recon_multiagent/
├── src/ Core deterministic layer (pipeline, agents, gates, DB, log)
├── config/ Externalized parameters (budgets, taxonomy, pricing)
├── evals/ Eval harness + metrics (recall, grounding, gaps, injection)
├── fixtures/ Cassette recordings + eval case definitions
├── migrations/ Postgres migrations (10 files, 8 tables)
├── scripts/ Smoke tests, cost script, secret gate, migration runner
├── tests/ Unit tests (Vitest)
├── workflows/ n8n workflow JSON exports, reference only
├── deploy/ Metabase docker-compose for on-demand ops
└── docs/ Runbook, dashboard SQL, demo plan, SRS/TDD, evidence
- SRS/TDD: controlled document, Rev 1.1 as built. The revision record covers the runtime decision, the async API, the HMAC scheme, and the eval framing.
- Runbook: redeploy, rotation, cassette refresh, DLQ, closeout.
- Dashboard build sheet: Metabase SQL.
- Demo plan: rate limiting and render.
- Config inventory: every tunable parameter.
This is Piece IV: orchestration. Bounded specialist agents, a hard-capped loop, and a deterministic gate between the model and the world.
- Piece I:
intake-n-outbound.pipeline. Pipeline automation. - Piece II:
document-intelligence-rag. Grounding and abstention. - Piece III:
shovels_n8n_nodes. Verified community nodes. - Piece IV:
recon_multiagent. You are here. - Capstone:
fieldops. Composes Recon's orchestration onto a real corpus with a human delivery gate.
Jake Morgan, jakemorganlabs
- Portfolio: jakemorganlabs.dev
- LinkedIn: linkedin.com/in/jakemorganlabs
- Contact: jakemorganlabs@gmail.com