Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
-
Updated
Aug 18, 2026 - Shell
Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
Copy inflation in multi-turn search agents: 78-92% of generated tokens are copied from retrieved documents and carry inflated log-probabilities, breaking confidence-based methods. Diagnostic toolkit + Retrieval-Grounded Voting. Findings of EMNLP 2026.
ATIF-native analytics over Claude Code agent trajectories: convert sessions to ATIF, materialize a corpus, query it with DuckDB.
VCR cassettes for agent trajectories: record agent runs as DAGs, replay the canonical path, only call the LLM for net-new paths
Local-first trajectory learning for coding agents. Every run makes the next one better.
Export Claude Code, Codex and Cursor sessions as ATIF trajectories. Rust core, Python CLI, installable with uv.
Open rubric and a small, hand-built gold set for judging multi-step web-research agent trajectories. The first dataset from Assayo.
Open rubric and a small, hand-built gold set for judging multi-step web-research agent trajectories. The first dataset from Assayo.
Independent research on trajectory-aware AI-agent evaluation, delegated authority, control integrity, and failure-preserving reproducibility.
Capture coding-agent sessions (Claude Code / Codex) at the source — trajectory + verifiable git environment — and score each session's training value (grounded × rich × focused). Local, deterministic, no model at runtime.
To associate your repository with the agent-trajectories topic, visit your repo's landing page and select "manage topics."