BenchFlow

Multi-turn agent benchmarking — Scene-based lifecycle for any ACP agent

What

BenchFlow runs AI agents against benchmark tasks in sandboxed environments. Single-agent, multi-agent, and multi-round patterns share one Scene-based lifecycle.

Any ACP agent — Gemini CLI, Claude Code, Codex, OpenCode, OpenClaw, Pi, or your own
Single + multi + progressive — single-agent / multi-agent (coder + reviewer, simulated user) / multi-round with a Python BaseUser callback
Cloud sandboxes — Daytona backend for parallel execution at scale
Hardened verifier — defaults block BenchJack/Meerkat-style reward-hacking; tasks opt out per-feature

Install

uv tool install benchflow

Requires Python 3.12+ and uv. Set DAYTONA_API_KEY for cloud sandboxes; export the relevant agent API key (GEMINI_API_KEY, ANTHROPIC_API_KEY, etc.) or run claude login / codex --login for subscription auth.

Documentation

Start with Getting started, then Concepts for the mental model. Then by goal:

If you want to…	Read
Run an eval on an existing task	Getting started
Understand Trial / Scene / Role / Verifier	Concepts
Author a new task	Task authoring
Multi-agent: coder + reviewer, simulated user, BYOS, stateful envs	Use cases
Multi-round single-agent (progressive disclosure, oracle access)	Progressive disclosure
Skill evaluation (when the artifact is a skill, not a workspace)	Skill eval
Understand the security model	Sandbox hardening
CLI flags + commands	CLI reference
Python API surface	Python API reference

Notebooks and runnable example scripts: examples/.

Featured

Progressive disclosure on SWE-bench Pro — the BaseUser abstraction drives a multi-round trial: terse round-0 prompt → failing-test hints → full spec. 5/5 oracle on Daytona, runnable demo at examples/swebench_pro_progressive_disclosure.ipynb. Also benchflow's Harbor #1316 parity answer for the no-second-LLM case. See Progressive disclosure.

Research artifacts

Two runnable labs validate the security story:

labs/benchjack-sandbox-hardening/ — end-to-end demo that 0.2.1+ blocks three BenchJack exploits that flip 0.2.0's reward from 0.0 to 1.0.
labs/reward-hack-matrix/ — full reward-hack sweep across real benchmarks comparing 0.2.0 vs 0.2.2.

Audience

Eval researchers / paper writers → Getting started → Concepts → Use cases
Task authors → Task authoring → Sandbox hardening
Agent builders integrating with benchflow → Concepts → Python API reference → benchflow.agents.registry
Existing Harbor users migrating → Use cases — migration section → Progressive disclosure (Harbor #1316 parity)

Contributing

PRs welcome. Open against main. CI runs ruff + tests on every PR; please run ruff check . and pytest tests/ locally first.

For a release: bump pyproject.toml to the next stable version, tag v<version> on main, push the tag — CI publishes to PyPI. Then bump main to the next .dev0.

License

Apache-2.0.

Name		Name	Last commit message	Last commit date
Latest commit History 615 Commits
.claude		.claude
.devcontainer		.devcontainer
.github/workflows		.github/workflows
benchmarks		benchmarks
docs		docs
examples		examples
experiments		experiments
labs		labs
src/benchflow		src/benchflow
tests		tests
.env.sample		.env.sample
.gitignore		.gitignore
.python-version		.python-version
CHANGELOG.md		CHANGELOG.md
CLAUDE.md		CLAUDE.md
LICENSE		LICENSE
README.md		README.md
pyproject.toml		pyproject.toml
uv.lock		uv.lock

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

BenchFlow

What

Install

Documentation

Featured

Research artifacts

Audience

Contributing

License

About

Uh oh!

Releases 7

Packages

Uh oh!

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

BenchFlow

What

Install

Documentation

Featured

Research artifacts

Audience

Contributing

License

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases 7

Packages 0

Uh oh!

Uh oh!

Contributors

Uh oh!

Languages

Packages