Can Language Models Solve Long-Horizon Software Security Tasks?
Hwiwon Lee, Jiawei Liu, Dongjun Kim, Wubing Xia, Ziqi Zhang, Chunqiu Steven Xia, Lingming Zhang
University of Illinois Urbana-Champaign
SEC-bench Pro is a repository for building advanced software security benchmarks from real-world bug reports, proof-of-concept inputs, and reproducible execution environments. The goal is to make difficult security cases easier to study, validate, and reuse for benchmarking research, triage workflows, and automated analysis systems.
The benchmark currently includes 344 verified cases across Chromium V8, Mozilla SpiderMonkey, and the Linux kernel. Each case packages meta.json, reproduction assets, Docker build files, validation logs, and fixed-image checks where applicable.
- [2026-06]: 🔥 SEC-bench Pro was used in the full GPT-5.5-Cyber and GPT-5.6 evaluation.
- [2026-06]: The Linux leaderboard is now live (link).
- [2026-05]: The SpiderMonkey leaderboard is now live (link).
- [2026-05]: SEC-bench Pro launched with the V8 leaderboard (link).
Prerequisites:
- Docker
- Python 3.11+
uv- Agent credentials for the config you run. Codex configs use
OPENAI_API_KEYor~/.codex/auth.jsonwithcopy_host_auth = true; Claude and OpenCode configs use the provider credentials declared in their TOML files.
Linux evaluations additionally require an x86-64 Linux host whose Docker daemon
can expose a readable and writable /dev/kvm to a privileged container. This
includes enabling nested virtualization when evaluating from a cloud VM.
The benchmark does not support QEMU TCG as an equivalent grading mode: it
changes execution time and scheduling, and can invalidate time-sensitive PoCs.
Check the actual container environment before starting an evaluation:
docker run --rm --privileged hwiwonlee/linux.x86_64:CVE-2022-0185 \
bash -lc 'test -r /dev/kvm && test -w /dev/kvm && echo KVM-ready'Install Python dependencies from the repo root:
uv syncRun an example Codex eval for a target project:
uv run harness/eval_codex.py harness/configs/codex/v8/config.example.toml
uv run harness/eval_codex.py harness/configs/codex/sm/config.example.toml
uv run harness/eval_codex.py harness/configs/codex/linux/config.example.tomlRun Claude or OpenCode with the matching config tree:
uv run harness/eval_claude.py harness/configs/claude/v8/config.example.toml
uv run harness/eval_opencode.py harness/configs/opencode/v8/config.example.tomlEach harness/configs/<agent>/<project>/ directory contains one
config.example.toml. Copy or edit that file to select another model,
provider, output directory, or instance set.
To evaluate a small target set, edit the TOML config before running:
instances = ["472139305"] # explicit case ids
# instances = "__verified__" # all verified cases in images_dir
outdir = "output/v8/codex/my-run"
images_dir = "../projects/v8" # use ../projects/sm or ../projects/linux for other targetsRelative outdir values are resolved under harness/, so the example above writes to harness/output/v8/codex/my-run/<timestamp>/<instance_id>/. See harness/README.md for all config fields, provider options, and artifact details.
Linux agent runs expose the KVM/QEMU harness through the vendored
mcps/linux package (secb-linux-vm-mcp). Agents write audit/poc.c and use
the trusted MCP tools secb_build, secb_repro, and secb_validate; the MCP
server runs outside the agent command sandbox so QEMU gets native KVM while the
agent shell remains network-restricted.
All Linux leaves now declare meta.json.privilege: user PoCs run as uid
1000, and root PoCs run as init-namespace uid 0. The Linux prompts and MCP
validator use this field so the stated attacker model matches the actual guest
execution.
The example configs harden network access per agent: Codex uses
workspace-write with network_access = false and web_search = "disabled",
Claude uses its Bash sandbox with denied domains, and OpenCode routes Bash tool
calls through a network-namespace shell wrapper.
harness/grade.py re-runs agent-produced PoCs against vulnerable, fixed, and latest images, then classifies each PoC with the project-specific judge prompt. Point --target-dir at one timestamped run or a parent directory containing timestamped runs:
uv run harness/grade.py --project v8 --target-dir harness/output/v8/codex/example/gpt-5.5 --benchmark-dir projects/v8 --pull-missing
uv run harness/grade.py --project sm --target-dir harness/output/sm/codex/example/gpt-5.5 --benchmark-dir projects/sm --pull-missing
uv run harness/grade.py --project linux --target-dir harness/output/linux/codex/example/gpt-5.5 --benchmark-dir projects/linux --pull-missingSummary CSVs are written to <timestamp_dir>/summary unless --out-dir is set. Linux latest validation uses per-CVE images from hwiwonlee/linux.x86_64.latest:<instance_id>; build local copies with:
python projects/linux/build_images.py --mode latest --linux-ref origin/master -j 4SpiderMonkey fixed images use hwiwonlee/sm.x86_64.fixed:<issue_id>. To build
a local replacement for a tag, run:
projects/sm/build_fixed_images.sh 1675905
# Or use a different repository and pass it to the grader with --fixed-repo.
projects/sm/build_fixed_images.sh --image-repo myorg/sm.x86_64.fixed 1675905Each project has a host-side oracle for checking the packaged PoC against its vulnerable image. These scripts require Docker and jq:
projects/v8/crash_check.sh 472139305
projects/sm/crash_check.sh 1880719
projects/linux/crash_check.sh CVE-2022-0185V8 and SpiderMonkey also accept a custom PoC path:
projects/v8/crash_check.sh 472139305 /path/to/poc.js
projects/sm/crash_check.sh 1880719 /path/to/poc.jsCheck that a fixed image mitigates the same case:
python projects/v8/patch_check.py 472139305
python projects/sm/patch_check.py 1880719
python projects/linux/patch_check.py CVE-2022-0185The crash oracles read image names, binaries, and command-line options from each case's meta.json. output.txt is retained as validation context, not as the scoring oracle.
base/
chromium/ Base image definitions for Chromium-related targets
linux/ Base/latest image definitions for Linux kernel benchmark cases
sm/ Base/latest image definitions for SpiderMonkey benchmark cases
v8/ Base image definitions for V8 benchmark cases
harness/
eval_codex.py
eval_claude.py
eval_opencode.py
common.py
grade.py
judge.py
configs/claude/
v8/
sm/
linux/
configs/codex/
v8/
sm/
linux/
configs/opencode/
v8/
sm/
linux/
prompts/
baseline/
judge/
mcps/
linux/ Vendored secb Linux VM MCP server
projects/
v8/
<issue_id>/ 103 Chromium V8 benchmark cases
sm/
<issue_id>/ 104 Mozilla SpiderMonkey benchmark cases
linux/
CVE-*/ 137 Linux kernel benchmark cases
build_images.py Build orchestrator for base/vuln/fixed/latest images
crash_check.sh Host-side vuln image crash oracle
patch_check.py Host-side fixed image mitigation oracle
This project is licensed under the MIT License. See LICENSE for details.
@article{lee2026sec,
author = {Lee, Hwiwon and Liu, Jiawei and Kim, Dongjun and Zhang, Ziqi and Xia, Chunqiu Steven and Zhang, Lingming},
journal = {arXiv preprint arXiv:2605.26548},
title = {{SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?}},
year = {2026}
}
@inproceedings{lee2025secbench,
author = {Hwiwon Lee and Ziqi Zhang and Hanxiao Lu and Lingming Zhang},
booktitle = {The Thirty-ninth Annual Conference on Neural Information Processing Systems},
title = {{SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks}},
url = {https://openreview.net/forum?id=QQhQIqons0},
year = {2025}
}