Trace Claude Code sessions in Braintrust via the Braintrust Claude Plugin
-
Updated
Sep 2, 2026 - Python
Trace Claude Code sessions in Braintrust via the Braintrust Claude Plugin
1st Place Winner (General Judge) - Datadog Self-Improving Agents Hack. Two identical AI agents play Split or Steal. No pre-programmed betrayal. They discover deception on their own. Built with @evancorrea.
The braintrust pi extension now lives in a new monorepo: https://github.com/braintrustdata/braintrust-coding-agent-plugins
My implementation for a kaggle competition: https://www.kaggle.com/competitions/WattBot2025
Trace Codex sessions in Braintrust via the Braintrust Codex Plugin
Braintrust Coding Agent Plugins
Can you eval an art form? Canon is a continuity linter for serialized TV, YouTube and micro-drama fiction. Canon plays the role of whats currently the scriptwriting coordinator, verifies your story's logic and cites the scene. No GenAI writing here. That's left to the humans.
Trace Antigravity sessions in Braintrust via the Braintrust Antigravity Plugin
Orchestrate other AI CLIs (agy, Codex, Grok, OpenCode, Claude Code) for second opinions, research, and codebase analysis. No Gemini CLI. Hybrid always-on skill (eval-backed).
Dual-modular platform combating SRE alert burnout and securing generative AI deployments using Spring Boot and FastAPI. Tracks on-call fairness via Gini coefficients and uses AST/CST analysis to detect fragile, low-quality, or overly AI-dependent code in CI/CD pipelines while introducing quantifiable engineering contribution metric
CacheCatch audits AI agent context and shows what to move so repeated tokens hit cache instead of full-price input — it finds context waste, prompt cache misses, and hidden agent cost leaks; and give you actionable insights how to fix them.
AI-powered coloring page generator for kids, parents, and teachers.
LLM evaluation platform — MMLU knowledge benchmarks across 57 subjects plus agentic evaluation with real finance tools, scored by an LLM judge.
Decision-grade comparison: LangSmith vs Langfuse vs Arize Phoenix vs Braintrust vs MLflow 3 for multi-team agent observability and evaluation. Versions, pricing and self-hosting terms verified against primary sources (August 2026).
A Python implementation of the canonical agent architecture: a while loop with tools. Build production ready AI agents with purpose-built tools, comprehensive tracing, and async patterns.
AI agent that turns regulated enterprise conversations into decision-ready briefs and audited actions. Permission-aware RAG with grounded citations, deterministic policy gates, HITL approval loop, and a 3-vertical eval scorecard.
A causal debugger for distributed AI systems—isolated Daytona worlds, Braintrust evals, and evidence-first incident response.
Bilingual LLM-as-a-judge benchmark and dashboard for measuring agreement, robustness, coverage, and when AI evaluators should abstain.
Darwin evolves the best whole agent for your task, its prompt, tools, code, and the model it runs on, generation over generation, inside sandboxes it can't escape and against a grader it can't game.
coding agent example solving for one-shot runnable applications
To associate your repository with the braintrust topic, visit your repo's landing page and select "manage topics."