[EMNLP 2026] MLOps-Bench: Benchmarking AI Agents on Production ML Engineering in Real Repositories
Figure 1. MLOps-Bench construction and evaluation pipeline. Source: companion manuscript.
MLOps-Bench evaluates coding agents on repository-grounded MLOps engineering changes. Each task provides a natural-language specification, a starting task workspace, a buggy baseline, an oracle implementation, and executable evaluation tests. Fail-to-pass (F2P) tests check newly required behavior, while pass-to-pass (P2P) tests protect existing behavior.
| Area | Release contents |
|---|---|
| Tasks | 212 executable benchmark tasks |
| Lifecycle coverage | Seven MLOps stage categories |
| Source provenance | 31 public source repositories with pinned commits and license material |
| Scoring tests | 636 F2P tests and 856 P2P tests |
| Execution model | Individual task workspaces or a compatible external harness |
The task distribution below comes from the dataset manifest.
| MLOps stage | Tasks | F2P tests | P2P tests |
|---|---|---|---|
| Data pipelines | 31 | 93 | 120 |
| Feature engineering | 30 | 90 | 118 |
| Model training | 29 | 87 | 106 |
| Model serving | 31 | 93 | 134 |
| Monitoring and observability | 31 | 93 | 130 |
| CI/CD and governance | 31 | 93 | 123 |
| End-to-end MLOps flow | 29 | 87 | 125 |
| Total | 212 | 636 | 856 |
- Git
- Python 3.9 or later
pytestto run the Python contract tests
No Oracle Cloud service, API key, build system, or deployment environment is required to inspect the dataset or run an individual oracle contract test.
Clone the repository and install the test runner in an isolated environment:
git clone https://github.com/oracle-samples/mlops-bench.git
cd mlops-bench
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip pytestTasks are organized by source repository, MLOps stage, and instance ID:
dataset/mlops-bench/<repository>/<stage>/<instance>/
Each task contains the following core material:
| Path | Purpose |
|---|---|
README.md and problem.json |
Task statement and requirements |
repo/ |
Starting task workspace |
buggy/ |
Intentionally incomplete baseline |
oracle/ |
Reference implementation and oracle contract test |
tests/test_f2p.py |
Tests for newly required behavior |
tests/test_p2p.py |
Tests for behavior that must remain stable |
metadata.json |
Task metadata, including published F2P/P2P counts |
expected_changes.patch |
Reference patch artifact |
To validate a selected task's oracle contract, run pytest from its
oracle/ directory. For example:
TASK=dataset/mlops-bench/01_GOOGLECLOUDPLATFORM__MLOPS_WITH_VERTEX_AI/01_data_pipeline/MLO-01_GOOGLECLOUDPLATFORM__MLOPS_WITH_VERTEX_AI-01_DATA_PIPELINE-20260728_194731
(cd "$TASK/oracle" && python -m pytest -q tests/test_stage_contract.py)For suite-wide agent evaluation, a compatible external harness must select a task, provide its workspace and specification to an agent, and apply the benchmark's evaluation tests. MLOps-Bench intentionally does not prescribe a single runner or agent integration.
F2P and P2P counts refer to public test_* functions in each task's
tests/test_f2p.py and tests/test_p2p.py files, respectively. The task-level
f2p_test_count and p2p_test_count metadata fields use the same convention.
The auxiliary oracle contract tests under oracle/tests/ are intentionally not
included in the published F2P/P2P totals.
Illustrative agent outcomes from the paper
Figure 2. Passing and failing agent outcomes reported in the companion manuscript. These examples illustrate behavioral evaluation, not a bundled user interface or runner.
The dataset manifest records the task and
test totals. The source-repository inventory
records public project URLs, pinned commits, license text locations, and
applicable upstream notices. The corresponding third-party license and notice
materials are available under
dataset/metadata/.
If you use MLOps-Bench, please cite the work:
@misc{tran2026mlopsbench,
title = {MLOps-Bench: Benchmarking AI Agents on Production ML Engineering in Real Repositories},
author = {Quoc Co Tran and Hitesh Laxmichand Patel and Diwakar Mahajan and Kshitij Bakliwal and Avi Sil and Katrin Kirchhoff},
year = {2026},
note = {Manuscript}
}This project welcomes contributions from the community. Before submitting a pull request, please review the contribution guide.
Please consult the security guide for responsible vulnerability disclosure.
Copyright (c) 2026 Oracle and/or its affiliates.
Released under the Apache License version 2.0.



