You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.
📚 Curated catalog of agent-training datasets + a toolkit that normalizes, deduplicates, and quality-tiers them into one schema. Produces 🤗 voidful/agent-sft.
Accepted to NeurIPS 2026 (Main Track). Reviewer precision and critique uptake are separable: a multi-agent protocol can detect errors well and still fail to repair them. Code, data and analysis for arXiv:2607.15388.
Unroll: read agent traces and chat logs in VS Code. Claude Code, Codex and other JSONL conversations as a readable chat. Public docs, examples and releases; implementation source is private.
Claude Code plugin that stops a coding agent from ending its turn with 'tests pass', 'build succeeds' or 'deployed' unless the transcript shows the command ran after the last edit and succeeded. Plus a CLI that audits agent transcripts for unsupported claims.
Retry Safety Analyzer: deterministic HIGH/MEDIUM/SAFE when a side-effecting tool may have been retried. Offline import — never executes the trace. HIGH is risk, not proof. External validation open.
Reference implementation, benchmarks, and traces for "Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers" (arXiv:2609.01466). Append-only trace ledger → deterministic run state → per-consumer views for the human observer and the agent.