> ## Documentation Index
> Fetch the complete documentation index at: https://docs.open4rena.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluations

> Run o4 on a set of coding tasks and compare how models do.

`o4 eval` runs o4 on a set of coding tasks, checks each result with a command
you choose, and records pass or fail, cost, and token use. Use it to compare
models, or prompt profiles, on tasks that look like your own work.

Evals call real models, so every run costs money. Set a budget before you run
a large set.

## How it works

An eval corpus is an `evals` folder in the directory where you run `o4 eval`.
It holds one TOML file per task. For each task and each model you pick, o4:

1. Copies the task's fixture (a small project) into a new temporary folder.
2. Runs o4 non-interactively on the task's prompt in that folder.
3. Runs the task's check command in the folder. The task passes if the command
   exits with code `0`.
4. Records the result in `.o4/evals.db` in the directory where you ran the
   command.

The model works only in the temporary copy, so your real files don't change.

## Set up a corpus

Create this layout in your project:

```text theme={null}
evals/
  corpus.toml
  tasks/
    fix-greeting.toml
  fixtures/
    greet/
      greet.sh
```

`evals/corpus.toml` holds a version string. Change it when you change the
tasks in a way that makes old results not comparable:

```toml theme={null}
version = "1"
```

Each file in `evals/tasks/` defines one task:

```toml theme={null}
[task]
id = "fix-greeting"
title = "Fix the greeting"
difficulty = "easy"
skills = ["edit"]
tags = ["smoke"]

[fixture]
kind = "dir"
path = "fixtures/greet"

[prompt]
text = "greet.sh prints the wrong greeting. Make it print exactly 'Hello, world'."

[check]
command = "sh greet.sh | grep -qx 'Hello, world'"
timeout_secs = 30

[limits]
timeout_secs = 300
max_turns = 20
token_cap = 50000
cost_cap_usd = 0.50
```

Check that o4 can read your tasks:

```text theme={null}
$ o4 eval list
ID	DIFFICULTY	SKILLS	TAGS	FIXTURE	SOURCE	SUBDIR
fix-greeting	easy	edit	smoke	dir	fixtures/greet	-
```

`--tag <TAG>` lists only the tasks with that tag, and `--json` prints the list
as JSON. If no task matches, the command stops with "no tasks matched the
requested eval selection".

If a task file is invalid, points to a `dir` fixture that doesn't exist, or
reuses an ID that an earlier file (in file name order) already has,
`o4 eval list` and `o4 eval run` skip it and print the
reason as a `warning: excluded eval task` line. o4 reads only `.toml` files
directly in `evals/tasks/`. A missing or invalid `evals/corpus.toml` is an
error.

### Task file reference

| Field | Required | Meaning |
| - | - | - |
| `task.id` | Yes | A unique ID. You select tasks by it with `--task`. |
| `task.title` | Yes | A short title |
| `task.difficulty` | Yes | `easy`, `medium`, or `hard` |
| `task.skills` | Yes | A list of labels, shown by `o4 eval list` |
| `task.tags` | Yes | A list of labels you can select with `--tag` |
| `fixture.kind` | Yes | `dir` or `git` |
| `fixture.path` | For `dir` | A folder, relative to `evals/`, copied for each run |
| `fixture.rev` | For `git` | A commit in the git repository you run `o4 eval` from |
| `fixture.subdir` | No | For `git`, a folder inside that commit to use as the workspace |
| `prompt.text` | Yes | The prompt o4 gets |
| `check.command` | Yes | A shell command run in the workspace after o4 finishes. Exit code `0` means pass. |
| `check.timeout_secs` | Yes | Seconds before the check is stopped |
| `limits.timeout_secs` | Yes | Seconds o4 gets to work on the task |
| `limits.max_turns` | Yes | The most turns o4 can take. At the limit, o4 gets one last turn to wrap up, and the check still runs. |
| `limits.token_cap` | Yes | The most tokens the task can use |
| `limits.cost_cap_usd` | Yes | The most the task can cost, in US dollars |

All four limits must be greater than zero.

A `git` fixture checks out the given commit of the repository in your current
directory into a temporary worktree, so tasks can start from a real past state
of your project.

<Note>
  The check command runs in o4's most restricted [sandbox](/safety/sandbox)
  tier, `airlock`: it can write only inside the task's folder and has no
  network access. A check that needs to download dependencies fails, so use
  offline commands such as `cargo test --offline`. On Linux the check needs
  bubblewrap; without it, every task is recorded as `skipped` with
  `fixture_error`.
</Note>

## Run evals

Pick the tasks and one or more models:

```bash theme={null}
o4 eval run --model anthropic/claude-sonnet-5 --tag smoke --budget 5
```

Use the same `provider/model` references as the rest of o4. See
[Choosing a model](/models/overview). o4 needs credentials for each provider,
the same as for a normal session. See
[Providers and API keys](/models/providers).

o4 prints a line as each task finishes and a summary at the end:

```text theme={null}
eval task=fix-greeting model=claude-sonnet-5 status=pass pass=1 fail=0 skip=0
eval run=1 pass=1 fail=0 skip=0 cost_usd=0.041200 status=completed
```

### Choose tasks

With no selection options, o4 runs every task in the corpus. These options
narrow it down, and a task must match all of the ones you use:

* `--tag <TAG>`: tasks that have this tag. Repeat it to allow several tags; a
  task needs any one of them.
* `--task <ID>`: a task by ID. Repeat it to pick several.
* `--difficulty <LEVEL>`: `easy`, `medium`, or `hard`.

If nothing matches, the command stops with "no tasks matched the requested
eval selection".

### Compare models

Repeat `--model` to run every selected task on each model:

```bash theme={null}
o4 eval run --model anthropic/claude-sonnet-5 --model openai/gpt-5.5 --tag smoke
```

If a model reference doesn't resolve, o4 warns and records that model's tasks
as skipped. If a model has no API key, its tasks are also recorded as skipped:

```text theme={null}
warning: skipping unresolvable model reference 'nope/nothing'
eval task=fix-greeting model=nope/nothing status=skipped pass=0 fail=0 skip=1
eval task=fix-greeting model=claude-sonnet-5 status=skipped pass=0 fail=0 skip=2
eval run=1 pass=0 fail=0 skip=2 cost_usd=0.000000 status=completed
```

If none of the `--model` references resolve, the command stops with
`could not resolve model reference` and records nothing.

### Limit cost

`--budget <USD>` stops the run once the total cost of finished tasks reaches
the amount. o4 checks the total after each task, so the last task can take the
run past the budget. A run stopped by the budget ends with `status=aborted`.
The budget can't be negative. Each task also has its own `limits.cost_cap_usd`.

Press `Ctrl+C` to stop a run. o4 records the task in progress as `aborted` and
marks the run as aborted.

### Other options

* `--prompt-profile <PROFILE>` uses one prompt profile for every task instead
  of the one o4 picks for each model. The values are `modern-minimal`,
  `modern-guided`, `legacy-guided`, and `local-defensive`. See
  [Reasoning and prompt profiles](/models/reasoning).
* `--enable-mcp` lets o4 use your configured MCP servers during the run. MCP
  is off in evals unless you pass it. See [MCP servers](/extend/mcp).

During an eval, o4 runs without approval prompts, and its shell commands are
sandboxed to the task's temporary folder. Your hooks and permission rules
don't apply.

## Read the results

`o4 eval report` shows the latest run as a table, with one row per model, a
column per task, and totals:

```text theme={null}
$ o4 eval report
MODEL           | fix-greeting | PASS | FAIL | SKIP | COST_USD
----------------+--------------+------+------+------+---------
claude-sonnet-5 | pass         |    1 |    0 |    0 |   0.0412
gpt-5.5         | fail         |    0 |    1 |    0 |   0.0388
```

* `--run <ID>` shows a different run. The run ID is in the summary line
  (`eval run=1 ...`).
* `--baseline <ID>` also lists the tasks whose result changed between the
  baseline run and the selected run. Both runs must use the same corpus
  version.
* `--json` prints JSON instead of a table.

A task run ends as `pass`, `fail`, or `skipped`, or `aborted` when you press
`Ctrl+C`. The JSON output gives the reason in `failure_category`, and a
`note` with details:

| `failure_category` | Status | Meaning |
| - | - | - |
| `check_failed` | `fail` | The check command exited with a non-zero code |
| `timeout` | `fail` | o4 ran out of `limits.timeout_secs` |
| `budget_exceeded` | `fail` | The task reached its `cost_cap_usd` or `token_cap` |
| `agent_error` | `fail` | o4 stopped with an error, such as a provider error |
| `fixture_error` | `skipped` | The fixture couldn't be set up, or the check didn't finish with an exit code (for example, it hit `check.timeout_secs`) |
| `missing_api_key` | `skipped` | o4 has no API key for the model's provider |
| `unresolved_model` | `skipped` | The `--model` reference didn't match a model |

`o4 eval summary` groups every recorded task run by the o4 version that ran
it, and shows how many were attempted and completed (passed), how many
completed with no intervention, and the matching rates. Skipped task runs
don't count as attempted; aborted ones do. Add `--json` for JSON output, which
also includes the definition of each measure.

## Compare prompt profiles

`o4 eval paired` runs the selected tasks twice per repeat: a control arm with
the prompt profile o4 picks for each model, and a candidate arm with the
`modern-minimal` profile. It prints `paired eval experiment=<ID>` first and
records each arm as its own run. It then compares the two arms on pass rate,
prompt tokens, and other measures, and prints a verdict for each model family:
`promote`, `reject`, or `insufficient-evidence`.

```bash theme={null}
o4 eval paired --model anthropic/claude-sonnet-5 --budget 10
```

It takes the same selection, `--budget`, and `--enable-mcp` options as
`o4 eval run`. `--budget` applies to each arm. `--repeats <COUNT>` sets how
many control and candidate pairs to run. The default is 3. `--json` prints
JSON.

With fewer than three complete repeats, the verdict can't be `promote`. If a
repeat stops early, for example on a budget or `Ctrl+C`, the command stops
with an error; the runs it recorded stay in `.o4/evals.db`. `--repeats 0` is
an error.

`o4 eval paired` runs two arms for every repeat, so with the default 3 repeats
it costs about six times as much as one `o4 eval run` of the same tasks.
