Skip to content

FEAT: Add Anthropic model-written evals dataset - #2986

Merged
Roman Lutz (romanlutz) merged 5 commits into
microsoft:mainfrom
Steeve-Crypto:feat/anthropic-model-written-evals
Oct 11, 2026
Merged

Roman Lutz (romanlutz) merged 5 commits into
microsoft:mainfrom
Steeve-Crypto:feat/anthropic-model-written-evals

Conversation

@Steeve-Crypto

Copy link
Copy Markdown
Contributor

Description

Adds the Anthropic model-written evals dataset as a remote seed dataset loader, so it can be fetched through SeedDatasetProvider.

The loader reads the full release from https://github.com/anthropics/evals. The Hugging Face dataset only publishes a subset, which is why this does not use load_dataset.

Supported collections are persona, sycophancy, advanced-ai-risk (human-written and LM-written eval files), and winogenerated examples. Few-shot generator prompts and the occupation catalog are left out because they are not eval items.

Rows with a question use that text as the seed prompt. answer_matching_behavior and answer_not_matching_behavior are kept in seed metadata, including when the non-matching field is a list. Winogenerated examples have no question column, so the prompt is sentence_with_blank and pronoun options stay in metadata.

This follows the current _RemoteDatasetLoader pattern. Closed PR #1170 targeted an older question-answering layout and was not merged.

Content warning: these prompts are meant to provoke model behavior and may contain offensive content.

Fixes #450

Tests and Documentation

Unit tests cover category filtering, behavior-label metadata, list-valued answers, winogenerated fill-in sentences, skipped non-eval files, GitHub listing errors, and an empty result. They use mocked responses and do not download the dataset.

Commands:

  • uv run pytest tests/unit/datasets/test_anthropic_model_written_evals_dataset.py (8 passed)
  • uv run pytest tests/unit/datasets/test_seed_dataset_provider.py -k _AnthropicModelWritten (metadata registration passed)
  • uv run ruff check and uv run ruff format --check on the changed files
  • uv run ty check pyrit/datasets
  • SKIP=ty-check uv run pre-commit run --files on the changed files

The pre-commit ty hook typechecks all of pyrit with --extra all and was not run. ty check on pyrit/datasets passed.

Jupytext was not run. The dataset is discovered through SeedDatasetProvider, so no notebook sample was added.

No CLA was signed from this change.

@Steeve-Crypto

Copy link
Copy Markdown
Contributor Author

@microsoft-github-policy-service agree

@romanlutz Roman Lutz (romanlutz) self-assigned this Oct 9, 2026
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Comment thread pyrit/datasets/seed_datasets/remote/anthropic_model_written_evals_dataset.py Outdated
Comment thread pyrit/datasets/seed_datasets/remote/anthropic_model_written_evals_dataset.py Outdated
Comment thread pyrit/datasets/seed_datasets/remote/anthropic_model_written_evals_dataset.py Outdated
Comment thread pyrit/datasets/seed_datasets/remote/anthropic_model_written_evals_dataset.py Outdated
@Steeve-Crypto

Copy link
Copy Markdown
Contributor Author

Roman Lutz (@romanlutz) Addressed all four review comments in 05d9fff. Ready for another look.

@romanlutz
Roman Lutz (romanlutz) added this pull request to the merge queue Oct 11, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 11, 2026
@romanlutz
Roman Lutz (romanlutz) added this pull request to the merge queue Oct 11, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 11, 2026
@Steeve-Crypto

Copy link
Copy Markdown
Contributor Author

Roman Lutz (@romanlutz) Thanks for the approval. The merge queue failed twice on test_sqlite_deadline_bounds_file_database_lock_wait_and_restores_timeout (Windows only, took ~1.2s vs the 1s limit). This PR doesn't touch memory code, so it looks flaky. Mind re-queueing when you get a chance?

@romanlutz

Copy link
Copy Markdown
Contributor

I'm fixing that in #3121 hopefully but need approval myself...

@romanlutz
Roman Lutz (romanlutz) added this pull request to the merge queue Oct 11, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 11, 2026
@romanlutz
Roman Lutz (romanlutz) added this pull request to the merge queue Oct 11, 2026
Merged via the queue into microsoft:main with commit 1bf811f Oct 11, 2026
55 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

FEAT Add Anthropic/model-written-evals Dataset

2 participants