Skip to content

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

1,572 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Harness (Behavior Layer for AI Coding Agents)

License

Harness is a lightweight, local behavior and orchestration runtime that wraps around your AI development sessions (Claude Code, Cursor, GitHub Copilot agent surfaces, Codex, Continue.dev, Hermes Agent). It provides reactive hooks, routing boundaries, and circuit breakers designed to prevent infinite trial-and-error loops, costly over-engineering, and "lost-in-the-middle" context drift.


The Problem

AI coding agents are highly capable, but they struggle with self-regulation, environment awareness, and attention limits:

  1. The Infinite Retry Loop: When an agent encounters a subtle compilation or test failure, its default behavior is to make micro-adjustments repeatedly (tweak and run, tweak and run) until it exhausts your token budget.
  2. Environment Blindness: Agents often assume standard Unix environments, hallucinating shell commands and paths when running on Windows, PowerShell, or sandboxed environments.
  3. Lost-in-the-Middle Bloat: As sessions grow, agents aggressively read too many large files or generate massive console logs, causing severe context degradation and reasoning hallucinations.

Why Harness?

Harness acts as an automated system supervisor. It remains completely silent and out of the way, intervening only when execution boundaries are violated or failures are detected.

Harness deliberately follows a minimal rails, maximum freedom design: semantic obligations are explicit, while runtime observation remains lightweight and mostly fail-open. It does not hard-stop work on iteration/revision/replan/worker counts. The one cognitive hard boundary is Rule of 3: three matching failures pause mutation for a zoom-out reflection. Router suggestions are conditional by applicability, not optional by default: the agent MUST read/evaluate each suggested skill; if applicable, its core contract MUST be followed, and if not applicable a flow-grounded reason MUST be retained. Implementation tactics remain flexible inside those obligations.

Comparison: Prompt vs. Skill vs. Harness

Dimension Prompt-Only (Custom Instructions) Skill-Only (Task Guides) Harness (Behavior Layer)
Activation Always loaded (wastes prompt space) Loaded on demand (requires manual trigger) Reacts dynamically when the selected host surface exposes compatible hooks/plugins
Fail-Safe No protection (model keeps retrying) No protection by itself Can add mechanical retry/verification boundaries on surfaces that package those mechanisms
Context Aware High risk of lost-in-the-middle bloat Manages scope manually Can add preflight/boundary mechanisms where the host exposes the required lifecycle/tool hooks
System Audit Blindly assumes shell syntax Requires manual shell check Uses environment detection and explicit verification paths instead of assuming one shell/runtime
Memory Resets on every new chat session Static text rules Stateful hook/plugin integrations can persist bounded runtime state; instruction-only integrations cannot

The exact "Harness" behavior depends on the installation surface. Claude Code has the broadest currently verified lifecycle-hook coverage. OpenCode has deterministic mechanism tests and partial live-host evidence for project-scope .js loading and edit/verification state on OpenCode 1.18.31 (macOS), not durable hard enforcement; global scope and npm install remain unverified. Codex has both an advisory --codex installer path and a local OpenAI plugin path with mechanism-tested session, prompt, supported-tool, subagent, and stop hooks. The public OpenAI Skills-only submission is narrower and does not include those local lifecycle hooks. See Supported AI IDEs & Tools and docs/platform-capabilities.md.

When should I use Harness?

  • You regularly use agentic coding tools (like Claude Code, Cursor, or Copilot) on medium-to-large codebases.
  • You develop on Windows or in mixed shells (Git Bash, WSL, PowerShell) where agents frequently get shell syntax wrong.
  • You want lightweight routing, objective verification, and failure-loop safety without forcing every task through a rigid workflow.

When should I NOT use Harness?

  • You only use chat interfaces for general questions without letting the AI run local commands or modify files.
  • You deliberately want completely unconstrained execution with no routing, verification, or retry boundaries.

⚡ Quick Start (Get Protected in 10s)

Harness integrates directly into your workspace. There is no heavy daemon, no paid external APIs, and zero configuration required.

Runtime: Harness supports Node.js 22+ (including odd-numbered versions); Node.js 24 is the primary development and CI runtime (.nvmrc). On Windows, System One evaluator subprocess verification should use Node 22 or Node 24.20+. Node 23.x has a documented Windows process-shutdown assertion (nodejs/node#56645); Node 24.0–24.19 predates the backported libuv fix (nodejs/node#61999) and may likewise crash at process exit (0xC0000409). RS13 visibly skips only on those Windows versions while retaining the real evaluator subprocess elsewhere. A skip is not a pass or evidence of a fixed Node runtime. CI pins Windows Node 23.6.0 and 24.15.0 to validate their skip policy, and Node 24.20.0 to run the complete subprocess/regression suite. These tests do not reduce the package's general >=22 engine compatibility claim. The generated current-state runtime/workflow summary is docs/repository-contract.md.

# Option A: Claude Code plugin (marketplace manifest included)
#   /plugin marketplace add dyphn1/Harness-everything
#   /plugin install harness-everything

# Option B: install Harness hooks/skills/advisory integrations into your workspace
npx github:dyphn1/Harness-everything install                 # POSIX shells / Git Bash
npx.cmd github:dyphn1/Harness-everything install             # Windows PowerShell

# Option C: install or update the native Claude Code/Codex plugin
#   npx github:dyphn1/Harness-everything plugin-sync       # POSIX shells / Git Bash
#   npx.cmd github:dyphn1/Harness-everything plugin-sync   # Windows PowerShell
#   ./scripts/plugin-sync.sh                               # POSIX shells / Git Bash
#   node scripts/plugin-sync.js                            # Windows-safe direct entry

# OpenAI/Codex local plugin packaging is repository-owned under:
#   .agents/plugins/marketplace.json
#   plugins/harness-everything/.codex-plugin/plugin.json
# See docs/openai-plugin.md for local import/install and public Skills-only submission.

Windows PowerShell: use npm-family .cmd shims such as npx.cmd rather than weakening ExecutionPolicy. When reading Harness Markdown from Windows PowerShell 5.1, specify UTF-8 explicitly, for example Get-Content -Encoding UTF8 -Raw <path>.

Expected Behavior After Installation:

  1. Use the selected surface's real mechanism: Claude Code hooks, the local OpenAI plugin lifecycle hooks, OpenCode's plugin API, or advisory instructions depending on what you installed.
  2. Preflight / session context where packaged: Hook-capable surfaces can inject environment/session context automatically; advisory-only surfaces must not be described as if they do.
  3. Verification boundary: Completion claims require objective evidence; whether that boundary is mechanically invoked or explicitly called depends on the host surface.
  4. Mandatory applicable workflow: MUST evaluate each suggested skill's complete SKILL.md; applicable core contracts MUST be followed, while not-applicable needs a flow-grounded reason. Selected-topology required obligations MUST resolve with objective evidence. Reasoning and implementation remain flexible; escape applies only to declared uncovered scope with evidence. Tier-3/Fable broad mutation MUST resolve isolation: verified linked worktree or explicit degraded fallback. See the workflow runtime contract. Harness does not impose one universal TODO/TDD/Fable sequence.
  5. Unified user-visible status: For non-trivial software/project work, the agent MUST render one scannable Markdown ### 🚦 Harness Status block with bold bullet labels for Current, Read / Evidence, Next, plus optional Risk / Blocked; multiple evidence items may use nested bullets. Use it at major phase/direction boundaries and before final completion. The routing checkpoint remains internal source state. This is a communication contract, not a hard runtime lock.

What Gets Installed (and How to Remove It)

The general installer only writes to your workspace (or, with --global, your home directory) — no daemons, no registry entries, no network services. Depending on which platforms you select, it creates:

File / Directory Purpose
.claude/settings.json (merged) + .claude/skills/ + .claude/agents/ Claude Code lifecycle hooks, project skills, and named Fable agents
.cursorrules + .cursor/skills/ Cursor advisory rules and project skills
.github/copilot-instructions.md + .github/skills/ GitHub Copilot agent-surface instructions and project skills
AGENTS.md + .agents/skills/ Codex advisory instructions plus repo-scoped Agent Skills; Hermes can also consume trusted project skills from .agents/skills/
.continue/rules/harness.md + .continue/skills/ Continue.dev advisory rule and installer candidate skill path; standalone skill discovery is Unknown
.hermes.md Hermes Agent project advisory context
.opencode/plugins/harness-enforcement.js (manual copy, not installer-owned) OpenCode enforcement plugin — the general installer has no --opencode path; copy opencode-plugin/index.mjs under a .js name per opencode-plugin/README.md
.claude/harness-everything/ (or the per-platform equivalent) Harness installer/runtime bookkeeping owned by that integration

For --global, the installer uses each host's supported user-level skill location rather than assuming one shared directory works everywhere: shared Agent Skills remain under ~/.agents/skills/ where natively consumed, Continue uses ~/.continue/skills/, Hermes uses ~/.hermes/skills/, and Claude uses ~/.claude/skills/.

The repository also ships a separate local OpenAI/Codex plugin package under plugins/harness-everything/ with marketplace metadata in .agents/plugins/marketplace.json. That package is not the same thing as the --codex advisory installer path. The public OpenAI Skills-only upload is narrower again; see docs/openai-plugin.md.

The native plugin-sync command detects the installed host CLIs and applies the state-specific operation: an absent plugin is installed, while an already-installed plugin is explicitly updated (Claude Code) or its configured marketplace is upgraded (Codex). Options the host CLI does not implement are dropped and retried, and a host build without plugin install commands is reported as an actionable skip. If a host cannot report plugin state, the command fails closed without installing or updating blindly.

The installer records its state directories in .git/info/exclude — a local-only git ignore file — so Harness state never lands in a commit and your working tree (including .gitignore) is never modified. Everything owned by the general installer is removed with the built-in uninstaller:

# POSIX shells / Git Bash
npx github:dyphn1/Harness-everything uninstall
npx github:dyphn1/Harness-everything uninstall --local --skills -y
npx github:dyphn1/Harness-everything uninstall --global

# Windows PowerShell
npx.cmd github:dyphn1/Harness-everything uninstall
npx.cmd github:dyphn1/Harness-everything uninstall --local --skills -y
npx.cmd github:dyphn1/Harness-everything uninstall --global

Visualizing the Flow

Without Harness (Endless Trial-and-Error Loop)

flowchart TD
    U([User Request]) --> A[AI Coding Agent]
    A -->|Command/Edit| Env[Workspace Environment]
    Env -->|Error / Failure| A
    A -->|Tweak & Retry 1| Env
    Env -->|Error / Failure| A
    A -->|Tweak & Retry 2| Env
    Env -->|Error / Failure| A
    A -->|Tweak & Retry 3... N| Env
    style A fill:#ffcdd2,stroke:#c62828,stroke-width:1px,color:#000000
Loading

With Harness (Invariant-First, Agent-Orchestrated Execution)

flowchart TD
    U([User Request]) --> K[Harness Kernel<br/>classify scope + establish invariants]
    K --> T{Tier classification}
    T --> S{Suggested skills?}
    S -->|Yes| R[Read each suggested SKILL.md<br/>evaluate flow + applicability]
    S -->|No| A[Agent chooses smallest useful tactic / skill set]
    R --> A
    A --> Exec[Execute Code / Run Commands]
    Exec --> Gate{Objective evidence supports completion?}
    Gate -->|No| Retry[Diagnose / iterate]
    Retry --> CB{Same-signature failure x3?}
    CB -->|No| Exec
    CB -->|Yes| ZO[Zoom Out / Re-plan]
    ZO --> Exec
    Gate -->|Yes| Done[Evidence-backed completion]
    Done --> SE[Optional Self-Evolve / Record]
    style K fill:#c8e6c9,stroke:#2e7d32,stroke-width:1px,color:#000000
    style CB fill:#fff9c4,stroke:#fbc02d,stroke-width:1px,color:#000000
    style ZO fill:#ffcc80,stroke:#ef6c00,stroke-width:1px,color:#000000
    style Gate fill:#ffcdd2,stroke:#c62828,stroke-width:1px,color:#000000
Loading

The Tier changes task shape and the set of suggested skills, not a universal required order. Tier 2 may suggest tdd, todo-driven-workflow, or verification-loop; Tier 3 may select Fable or multi-agent topologies. Every suggestion is MUST-evaluate: if its real flow is applicable, the core contract becomes MUST-follow; otherwise retain a flow-grounded not-applicable reason. Tactics inside the resulting contract remain model-controlled.


Core Modules & Concepts

Harness operates through six core cognitive concepts:

  1. Kernel Router (kernel-router.js + tier-router.js): tier-router.js remains the classifier, dynamic-skill detector, and knowledge-guide matcher. kernel-router.js is the public runtime boundary: it preserves the classifier result, injects baseline MUST invariants, and adds evaluate-suggestions-before-skip whenever domain skills are suggested. Suggestions MUST be evaluated from their real SKILL.md flow; an applicable skill's core contract MUST be followed, while not-applicable needs evidence. Selected-topology obligations are also semantic MUSTs, but the runtime generally observes/reminds rather than hard-blocking. This preserves agent autonomy over tactics without allowing confidence to erase the lifecycle. If nothing matches at all — including nothing already kept from the open skills ecosystem — find-skills checks npx skills list live and, if still nothing, searches skills.sh/npx skills with explicit approval before installation.
  2. Guard (rule-of-3.js): The fail-safe circuit breaker. Tracks failure signatures across terminal runs on integration surfaces that package the required lifecycle hooks. If a test or command fails 3 times with the same signature, it locks mutating tools and forces a zoom-out reflection: re-verify every assumption with read-only tools, write a fact-checked report, then resume on a fresh diagnosis. A companion Stop hook (stop-gate.js) emits a non-blocking reminder when edits were never followed by successful verification on hosts where that hook is installed.
  3. Memory (state-persist.js): Session transaction logging for stateful hook/plugin integrations. Static skills/instructions alone do not create WAL state.
  4. Reflection (self-evolve): Long-term workspace immunization. Upon task completion, the agent reflects on the root cause of resolved issues, then judges whether the lesson is a simple rule or a reusable, complex pattern: simple rules are appended to local workspace rules (RULES.md); genuinely reusable patterns are instead packaged as a dynamic skill (via skill-creator's Dynamic Skill Generation Contract) and registered in manifest.json so the Router picks it up in future sessions. Either path is validated by a hermetic self-regression suite before it's persisted.
  5. Subagent Scope Guard (subagent-scope-guard.js): Diffs the whole repo's git status before and after every supported subagent (Task) burst, not just the files it was briefed to touch. Catches a subagent that was told to only read/verify but edited files anyway — where that host/integration actually invokes the guard.
  6. Cognitive Laws (Agent Cognitive OS): The Cognitive OS is a policy layer, not a peer skill that must win host routing before domain work can begin. Its Discover → Think → Try → Summarize → Record loop remains available as an explicit/manual entry point, while runtime integrations establish the smaller cross-cutting invariants independently.

Supported AI IDEs & Tools

The authoritative current matrix is docs/platform-capabilities.md. The important distinction is that one host can have multiple Harness installation surfaces.

opencode-plugin/ (index.mjs) implements verification observation/reminders and a Rule of 3 breaker against OpenCode's real, source-verified plugin API (tool.execute.before/.after, the session.idle event). ci/mechanism-2n-opencode-plugin.test.js drives the exported hooks directly against a mock context matching that API. Retained evidence supports project-scope .js plugin loading and edit/verification state on OpenCode 1.18.31 (macOS): live-host evidence. The retained snapshot predates #190's simplification and is historical evidence only: its final snapshot was post-reset; the old hard-lock was only an interactive observation, with no retained blocked-tool trace. Current OpenCode behavior keeps the third-failure reflection boundary but no permanent post-reflection hard lock. Reflection was operator-seeded, then agent-rewritten, not unaided; no behavioral-effectiveness claim is made. Install with a .js destination name: this host silently ignores .mjs (issue #127, guarded by test:opencode:loadability). Global scope, npm-package installation, and other host versions remain unverified.

AI Agent Tool / Surface Integration Method Local Target Location Enforcement claim
Claude Code Native lifecycle hooks (PreToolUse, PostToolUse, SessionStart, UserPromptSubmit, Stop) .claude/settings.json, .claude/skills/, .claude/agents/ Semantic workflow contracts are reminder-observed; Rule-of-3 reflection and explicit permission boundaries may block
OpenCode Native plugin module (opencode-plugin/index.mjs, install as harness-enforcement.js) .opencode/plugins/ Partial live-host evidence — project-scope .js loading and edit/verification state on OpenCode 1.18.31 (macOS); current workflow/verification behavior is reminder-oriented, with only the third-failure reflection boundary blocking edits
Codex — general installer path Skills + AGENTS.md instructions AGENTS.md + repo-scoped .agents/skills/ Instruction/advisory delivery only; semantic MUST/SHOULD/MAY still applies, without a hard-enforcement claim
Codex / local OpenAI plugin .codex-plugin package with session, prompt, supported-tool, subagent, and stop hooks plus 26 canonical skills .agents/plugins/marketplace.json → plugins/harness-everything/ Mechanism-tested local contract observation plus explicit permission/Rule-of-3 boundaries for the packaged mappings; live host loading remains unverified
Public OpenAI Skills-only plugin Public Skills-only bundle Generated submission ZIP from plugins/harness-everything/skills/ Skill/workflow behavior only; no local .codex-plugin lifecycle hooks in the public artifact
Cursor Native Project Rules + project skills .cursorrules + .cursor/skills/ Advisory only
GitHub Copilot agent surfaces Agent Skills + repository custom instructions .github/copilot-instructions.md + .github/skills/ Agent Skills path is documented; no live Harness session or plugin install is verified
Continue.dev Native project rules; skill path retained as an installer adapter candidate .continue/rules/harness.md + .continue/skills/; global candidate ~/.continue/skills/ Rules are documented; standalone SKILL.md discovery is Unknown
Hermes Agent Trusted project context + skills .hermes.md + trusted project .agents/skills/; global skills ~/.hermes/skills/ Skill path and installer contract are checked; project loading remains subject to trust

For local OpenAI packaging, marketplace import, plugin tests, and the public Skills-only submission boundary, see docs/openai-plugin.md.


Repository Index

Multi-Agent Workspace

Use the canonical multi-agent-workspace skill for permanent multi-agent infrastructure:

node multi-agent-workspace/scripts/scaffold.js --workspace . \
  --agency-source <path-to-agency-agents> --division engineering --platform codex

The source is read-only input. Runtime metadata, selected roles, the launcher, resolved router, memory index, and structured handoff are keyed under the global Harness state home; no generated router, executable, or zone skeleton is written to the target workspace. Decision, domain, and architecture records are resolved per repository from CONTEXT-MAP.md, project configuration, existing documentation folders, or a committable fallback. Omit the source for an explicit unavailable-catalog fallback; do not treat it as a complete roster.

This repo uses a flat layout (waza/agentskills.io convention). The table below maps each top-level directory to its role.

Directory Category Description
harness-everything Core Runtime Bootstrap, kernel-router, tier-router, verify-gate, self-heal
hooks Core Runtime Claude Code lifecycle hooks (prompt routing, circuit breaker, scope guard, stop gate, etc.)
scripts Core Runtime Installer, manifest, prompts, workspace utilities, repository contract extraction/sync
bin Core Runtime harness CLI entry point
ci Quality Gates Consistency checks, description collision, mechanism tests, invariant-routing regression, documentation/runtime contract drift guards
.github CI/CD GitHub Actions workflows (ci.yml, release.yml, behavioral-evals.yml)
.claude-plugin Distribution Plugin manifests for Claude Code marketplace
.agents/plugins Distribution OpenAI/Codex repository marketplace metadata
plugins/harness-everything Distribution Local OpenAI/Codex plugin package plus canonical skill copies
submission/openai Distribution / Review Public OpenAI Skills-only listing/test inputs
evals Routing Evals 26 trigger/routing eval suites (waza format)
eval-framework Quality Gates Negative-control fixtures for consistency/collision gates (not a skill)
contract-integrity Quality Gates ADR→spec→ticket→test→implementation trace audit (Phase 1; not a directly routed skill)
telemetry Quality Gates Local JSONL operational evidence layer with report/benchmark scripts (not a skill)
behavioral-evals Behavioral Evals LLM-level discipline cases plus weekly structural validation workflow
benchmarks Benchmarks BENCHMARK_SOP fixtures and recorded A/B results
docs Documentation Philosophy, architecture, routing, reflection, platform capabilities, generated repository contract, audit
references Documentation Shared checklists (security, performance, definition-of-done)
multi-agent-workspace Skill (Tier 3) Scaffold a verified multi-agent workspace and select bounded specialists from an external catalog without vendoring the full roster
environment-detection Foundation Preflight: detect OS, shell, package manager
eval-harness Skill (Tier 2) Evaluate agent outputs against rubrics
fable-discipline Skill (Tier 3) Fable execution guardrails when Fable is selected
fable-mode Skill (Tier 3) Optional macro/multi-agent orchestration with milestone gates
find-skills Meta Discover and install skills from open ecosystems
git-commit Skill (Tier 1) Conventional commit messages with verification
grill-me Skill (Tier 2) Adversarial plan interrogation before implementation
grill-with-docs Skill (Tier 3) Domain-model and decision alignment before design publication
improve-codebase-architecture Skill (Tier 2) Architectural refactoring with evidence
install-cognitive-os Foundation / Manual Entry Explain or explicitly apply the cognitive policy; runtime invariants do not depend on host selecting it
repo-docs Skill (Tier 3) Generate repository documentation
rewrite-commits Skill (Tier 1) Interactive rebase and commit history cleanup
security-review Skill (Tier 2) OWASP/STRIDE security review
self-evolve Skill (Tier 2) Workspace immunization via dynamic skills
skill-creator Meta Create new skills from patterns
skill-style Meta Skill authoring style guide
tdd Skill (Tier 2) Test-driven development when executable behavior benefits from it
to-spec Optional (MAY, Tier 2/3) Publish specs from settled conversations
to-tickets Optional (MAY, Tier 2/3) Decompose settled specs into tracked tickets
todo-driven-workflow Optional Foundation (MAY) Progress tracking when explicit multi-step state helps
using-git-worktrees Skill (Tier 2) Git worktree concurrency patterns
verification-loop Skill (Tier 2) Select systematic verification evidence; kernel still requires evidence before completion
verify-before-claim Always-on discipline Fact-audit before asserting claims
zoom-out Circuit breaker Circuit-breaker reflection protocol
opencode-plugin Platform Plugin Enforcement logic for OpenCode's real plugin API; live-session firing still unverified (#37)

Deeper Documentation

For a deep dive into individual modules and the underlying philosophy, explore our sub-documents:

Fable model selection is documented in fable-mode/references/model-matrix.md; the explicit entrypoints are fable-haiku, fable-sonnet, and fable-opus.

Maintainers MUST follow RELEASING.md for tag-driven npm releases and record observations in docs/release-evidence.md. Issue #20 is closed (2026-09-10); its coordination history lives in git.


Experimental: System One routing (off by default — keep it off for now)

Issue #233 adds an optional local "System One" scorer. It scores the fixed tier catalog for each prompt. The lexical router stays authoritative:

  • HARNESS_SYSTEM_ONE_MODE defaults to off.
  • shadow only logs a diagnostic.
  • prefer can raise a tier but never lower it below the deterministic structural floor, and only for a harness-routing-v1 checkpoint.

Honest status: it is not useful yet, so do not enable it for normal work.

  • No usable model yet. The only available checkpoint, cua-ai/cua-s1-forms, is a form-filling model, so every decision abstains as domain-mismatch and routing never changes. On historical routing keywords its top-1 skill choice is at chance level: 8.2% against 11.6% for a random guess, and 5.9% against 3.8% on eval prompts. It is kept only to prove the pipeline. The feature should stay off until a Harness-trained checkpoint passes the Phase 4 gates.
  • It costs something when on. The resident transport keeps one Python/PyTorch server per manifest alive until 30 minutes idle. The measured cost on one Windows CPU host is about 85–110 ms per prompt, and the warm p95 of 108.6 ms misses the 100 ms target. The first prompt after a cold start uses the lexical route while the server loads.
  • Setup downloads a lot. npm run system-one:install needs Python 3.11–3.13 and downloads CPU PyTorch (~124 MB on Windows), the pinned cua_s1 source and the pinned checkpoint. It never runs from hooks and never turns the mode on.

Enabling it depends on the host, and most options are process-wide:

Surface How System One gets its configuration Caveat
Claude Code (plugin or --claude install) env in ~/.claude/settings.json (HARNESS_SYSTEM_ONE_MODE, HARNESS_SYSTEM_ONE_CONFIG) The values reach hooks and every command the agent runs (tests, builds), and are hot-reloaded into running sessions.
Codex (plugin or --codex install) Hooks inherit the environment of the codex process A Codex started from Claude Code inherits Claude's env. A Codex started from an ordinary terminal only sees the variables if they are set at user/OS level (Windows user environment variables, or your shell profile on macOS/Linux), which then applies to every program you start.
Skills-only / public OpenAI bundle No hooks It runs only if the agent itself runs harness-everything/scripts/kernel-router.js.

Evidence so far:

  • Mechanism: Linux, macOS and Windows CI.
  • Real CPU inference: on one Windows machine.
  • Live host: one Windows + Claude Code observation, where a hook-spawned server outlived the host. macOS/Linux hosts and Codex hosts are not live-verified.
  • Details: docs/system-one-routing.md and benchmarks/results/system-one/.

Benchmarks & Testing

If you are an agent asked to verify a Harness install, start at VERIFICATION.md, not here. It separates package/integrity checks, mechanism evidence, live-host evidence, and behavioral evidence so one kind of pass is not mistaken for another.

npm test (self-evolve/scripts/self-regression.js) runs deterministic syntax, CLI, routing-matrix, positive skill-route coverage, invariant-first routing regression, reference, behavioral-case, Fable model-mode, and mechanism checks (ci/mechanism-test.js, npm run test:mechanism to run it alone). The suite checks real exit codes and stderr, not just "the code looks right." Primary CI runs on Node.js 24 across Ubuntu, Windows, and macOS for every push and pull request (.github/workflows/ci.yml), with a separate Node.js 22 compatibility lane for the advertised minimum runtime.

For a fuller vanilla-vs-Harness behavioral comparison, see Harness Skills Benchmark SOP — standardized, reproducible scenarios:

  • Test A: Over-engineering defense (Tier 1 typo correction)
  • Test B: Micro-error loop defense (Tier 2 bug resolution)
  • Test C: Attention loss and hallucination (Tier 3 module refactoring)
  • Test D: Knowledge boundary constraints (Offline hallucination prevention)
  • Test E: Terminal environment and shell awareness (Windows/Unix shell detection)
  • Test F (in VERIFICATION.md, not BENCHMARK_SOP.md): fact-audit discipline — does the agent verify an external-behavior claim before asserting it?

Benchmark results are tracked in benchmarks/ (run.js scaffold builds the fixture, record commits a schema-validated result bound to a session log). Until those cells are filled, effectiveness claims are unbacked by recorded evidence.

Behavioral evals (LLM-level, on demand + weekly structural validation)

Mechanism tests prove individual packaged mechanisms; only real host/session evidence proves the host actually loaded and fired them. behavioral-evals/ runs discipline cases (including pressure variants like "we ship in 5 minutes, skip checks") against headless agent sessions (claude -p, OpenCode) in throwaway workspaces: npm run eval:behavioral. Live runs remain token-costing and can be invoked on demand. The weekly behavioral-evals.yml workflow always validates case structure and runs live cases only when the runner actually has the Claude CLI.

Catalog hygiene

npm run test:consistency keeps distribution manifests, docs links, skill frontmatter, routing-eval coverage, platform capability claims, and the generated repository runtime/workflow contract in lockstep with what is actually on disk. npm run test:repo-contract runs the runtime/workflow drift gate directly, while npm run docs:sync regenerates docs/repository-contract.md after an intentional change. npm run test:docs:capabilities runs the platform-doc drift check directly. npm run test:references checks every executable/deep-dive path named by SKILL.md; npm run test:release compares release/catalog evidence; npm run test:routing:skills recursively classifies nested skills and executes the real router for every directly-routable skill's positive cases; npm run test:routing:invariants guards the invariant-first architecture including read-before-skip; and harness verify-install detects stale installed versions or file trees. The installer E2E gate performs install → verify-install → uninstall against seeded user-owned files on Ubuntu, Windows, and macOS so path and ownership symmetry regressions fail CI. npm run test:collision fails CI when two skills' descriptions overlap enough to confuse the router.


📊 System Evaluation

Harness audits itself on a dated cycle by running its own test suite and VERIFICATION.md recipes — never by reading the code and assuming it works. The full scorecards, methodology, and per-cycle change log live in docs/audit.md. Audit scorecards are historical snapshots; current platform capability claims live in docs/platform-capabilities.md.

Latest local audit baseline — 2026-09-18 (macOS, Node.js 24; see docs/audit.md): 26/26 on-disk skills, 26/26 routing-eval directories, 34/34 positive routes, 275/275 invariant checks, and all 26 canonical skills within the 500-token limit. Deterministic gates were green except the pre-existing mechanism-30 failure caused by leftover .worktrees fixtures (reproduced on the clean tree). Waza full-matrix spec verify and live model sessions remain on-demand evidence; run waza on an LF-normalized export as documented by CI.

Measure on an LF export, not a Windows working tree — CRLF can inflate waza's token counts and trigger false budget failures.


🤝 For Contributors

To contribute to Harness or modify any Skill behavior, ensure you run the local self-regression suite and consistency gates first:

npm run self-regression
npm run test:consistency
npm run test:repo-contract
npm run test:plugin:openai
npm run test:plugin:submission

After intentional runtime/workflow changes, run npm run docs:sync and commit the regenerated repository contract. All script modifications must pass 100% cleanly before pushing to keep the runtime immunized against behavioral regression.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages