Skip to content

About

Piece 4 in the Field Ops portfolio. Multi-agent company intelligence. Bounded Search-Analyst-Synthesis loop. Deterministic grounding.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Recon

Three-agent company-intelligence system. Bounded loops, deterministic grounding, cassette evals, provenance-attested releases.

CI Release Eval Gate

Status

v1.1.0. Deployed and live on a single Hetzner VPS behind a Cloudflare tunnel. Zero open inbound ports.

  • Live demo: https://recon.jakemorganlabs.dev (HMAC-signed POST /webhook/recon; public GET /health; rate limit 5 req / 10 s / IP)
  • Sample dossier: docs/evidence/sample_dossier.md (a real live run, source-cited)
  • Eval report: docs/evidence/eval_report.md (30-case cassette suite, reproducible offline)

The runtime is a deterministic TypeScript orchestrator that systemd runs. It is not the n8n workflows in workflows/. See the architecture note below. No evidence file contains a real secret.

What it does

Recon turns a company name into a structured, source-cited intelligence dossier.

  1. One bounded extraction captures the brief.
  2. A Search agent gathers evidence from the open web.
  3. An Analyst agent extracts signals from that evidence.
  4. A Synthesis agent writes the dossier.
  5. The agents coordinate through inspectable shared state.

Accountable orchestration is the defining property. Each agent gets a minimal, role-specific toolset. A schema gate checks each output before the next agent reads it. Every claim traces to a retrieved evidence snippet. The system takes no autonomous action on the world. The worst case is a dossier with explicit gaps, never a fabricated claim and never an action taken.

Architecture

flowchart LR
    subgraph Brief["Brief Normalizer"]
        A[Raw Request] --> B[Structured Brief]
    end

    subgraph Search["Search Agent (bounded web)"]
        B --> C[Web Queries]
        C --> D[Page Fetches]
        D --> E[Evidence Store]
    end

    subgraph Analyst["Analyst Agent (no web)"]
        E --> F[Signal Extraction]
        F --> G{Coverage Check}
        G -->|Unfilled slots & budget| C
        G -->|Done| H[Signals]
    end

    subgraph Synthesis["Synthesis Agent (no web)"]
        H --> I[Dossier Draft]
        I --> J[Grounding Gate]
    end

    subgraph Output["Output"]
        J -->|Pass| K[Grounded Dossier]
        J -->|Fail| L[Gapped Dossier]
    end
Loading

The no-action posture is the safety property. The system only produces a dossier a human reads. It sends nothing, writes nothing, transacts nothing. The worst outcome of any agent error is a report that flags what it could not verify.

A note on the runtime (deviation from the 1.0 baseline)

The 1.0 SRS/TDD specified the three agents as n8n AI Agent nodes, with the n8n workflow graph as the orchestrator. The shipped and deployed system is a deterministic TypeScript orchestrator (src/pipeline.ts). The agents, the bounded coverage loop, the grounding gate, and the schema checks are all TypeScript, run as a systemd service behind an HMAC endpoint. This was a deliberate change: cohesion with the other portfolio pieces, and deterministic, testable control flow. The workflows/ directory retains n8n JSON exports as reference triggers only. They are not the running system, and Recon does not need them. The TypeScript core is the source of truth. Rev 1.1 of the SRS/TDD records this decision and its reasons.

The request pattern

A full run outlives the tunnel's edge timeout. The API is therefore asynchronous.

  1. Send a signed POST /webhook/recon. The service verifies the signature, creates the run, and returns 202 with a run_id.
  2. Poll GET /run/<id> until the status is terminal.
  3. Read the dossier from the terminal response.

NOTE: An identical request body derives the same idempotency key and returns the existing run. Vary the target to force a fresh run.

The measured bar

Recon ships with a labeled eval set of 30 cases: 10 rich, 8 thin, 5 empty, 7 adversarial. The suite runs against recorded cassette fixtures. No live web is required. The suite gates CI and releases.

Suite Cases Metric Threshold Gate
Evidence Recall (recall@k) 10 rich >= 0.80 1.00 Hard
Structural Validity 30 all >= 0.95 1.00 Hard
Grounding Integrity (recast-gap rate) 30 all <= 0.05 0.0000 Hard
Gap Correctness (FAR / FAR_inv) 10 thin <= 0.05 / <= 0.20 0.00 / 0.01 Hard
Injection Resistance (instructions obeyed) 7 adversarial 0 0 Hard

Result: 30 of 30 cases pass. All gates are green. These are deterministic-replay numbers. The suite runs offline against recorded web cassettes with a deterministic LLM stub, so it is a reproducible regression gate, not a measure of live model quality. Recall and validity at 1.00 partly reflect internally consistent fixtures. The meaningful hardening result is injection resistance: 0 of 7 adversarial instructions obeyed. That metric exercises real control flow. The agents route injected page text as data, never as instructions. Live-quality evidence is the sample dossier in docs/evidence/, produced by a real signed run over the tunnel. The thresholds shown are the gates that evals/run.ts enforces. Regenerate with npm run eval (needs Postgres, no web, DEEPINFRA_API_KEY unset).

Security posture

Recon's security is structural, not just procedural.

  • No-action posture. The worst case is a dossier with explicit gaps. The system never sends an email, writes an external row, or calls an API outside the evidence store.
  • HMAC-signed webhooks. Every inbound request is signed with a shared secret. A stale request (over 5 minutes) or an invalid signature is rejected.
  • Secrets only in an on-box env file. deploy/.env.production is chmod 600, gitignored, and loaded by the systemd service. It never enters git and never leaves the box. scripts/secret_gate.sh greps the tree for token patterns and fails CI on any literal-key hit.
  • OIDC-attested releases. Every release artifact carries a GitHub-signed provenance attestation via actions/attest-build-provenance.

Observability & economics

Dashboard: a Metabase compose file ships in deploy/metabase/docker-compose.yml but is not deployed. The single 1.9 GB VPS already runs three Node services plus host Postgres. An always-on Java dashboard is the largest OOM risk on that box, so it stays an on-demand operator tool, not a live service. Run-level observability lives in Postgres: the audit, tool_call, and run tables log every stage transition, evidence fetch, and status change per run_id.

make cost-month
Metric Value Source
Cost per run (p50) not yet metered scripts/cost_monthly.sh
Cost per run (p95) not yet metered scripts/cost_monthly.sh
Cache savings rate 0 (no prompt caching on DeepInfra) scripts/cost_monthly.sh

Gemma 4 on DeepInfra does not support prompt caching, so the cache columns stay zero. The metric structure is in place for future model swaps.

Run it

Local quickstart

git clone https://github.com/jakemorganlabs/recon_multiagent.git
cd recon_multiagent

cp .env.example .env
# set your DeepInfra key and search API key for live runs
npm install

docker run -d -e POSTGRES_USER=recon -e POSTGRES_PASSWORD=recon -e POSTGRES_DB=recon -p 5432:5432 postgres:18
npm run migrate

npm test

npm run eval   # offline cassette suite; leave DEEPINFRA_API_KEY unset for deterministic replay

WARNING: Do not run the eval with a live DEEPINFRA_API_KEY set. A live model rewords the recorded queries, every cassette lookup misses, and the numbers are meaningless. The deterministic mode activates only when the key is unset and CASSETTE_MODE=play.

Production

See docs/runbook.md for redeploy, secret rotation, cassette refresh, DLQ checks, and the closeout protocol.

Repo map

recon_multiagent/
├── src/            Core deterministic layer (pipeline, agents, gates, DB, log)
├── config/         Externalized parameters (budgets, taxonomy, pricing)
├── evals/          Eval harness + metrics (recall, grounding, gaps, injection)
├── fixtures/       Cassette recordings + eval case definitions
├── migrations/     Postgres migrations (10 files, 8 tables)
├── scripts/        Smoke tests, cost script, secret gate, migration runner
├── tests/          Unit tests (Vitest)
├── workflows/      n8n workflow JSON exports, reference only
├── deploy/         Metabase docker-compose for on-demand ops
└── docs/           Runbook, dashboard SQL, demo plan, SRS/TDD, evidence

Docs index

  • SRS/TDD: controlled document, Rev 1.1 as built. The revision record covers the runtime decision, the async API, the HMAC scheme, and the eval framing.
  • Runbook: redeploy, rotation, cassette refresh, DLQ, closeout.
  • Dashboard build sheet: Metabase SQL.
  • Demo plan: rate limiting and render.
  • Config inventory: every tunable parameter.

Part of a five-piece portfolio

This is Piece IV: orchestration. Bounded specialist agents, a hard-capped loop, and a deterministic gate between the model and the world.

  • Piece I: intake-n-outbound.pipeline. Pipeline automation.
  • Piece II: document-intelligence-rag. Grounding and abstention.
  • Piece III: shovels_n8n_nodes. Verified community nodes.
  • Piece IV: recon_multiagent. You are here.
  • Capstone: fieldops. Composes Recon's orchestration onto a real corpus with a human delivery gate.

Author

Jake Morgan, jakemorganlabs

About

Piece 4 in the Field Ops portfolio. Multi-agent company intelligence. Bounded Search-Analyst-Synthesis loop. Deterministic grounding.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages