o4 eval runs o4 on a set of coding tasks, checks each result with a command
you choose, and records pass or fail, cost, and token use. Use it to compare
models, or prompt profiles, on tasks that look like your own work.
Evals call real models, so every run costs money. Set a budget before you run
a large set.
How it works
An eval corpus is anevals folder in the directory where you run o4 eval.
It holds one TOML file per task. For each task and each model you pick, o4:
- Copies the task’s fixture (a small project) into a new temporary folder.
- Runs o4 non-interactively on the task’s prompt in that folder.
- Runs the task’s check command in the folder. The task passes if the command
exits with code
0. - Records the result in
.o4/evals.dbin the directory where you ran the command.
Set up a corpus
Create this layout in your project:evals/corpus.toml holds a version string. Change it when you change the
tasks in a way that makes old results not comparable:
evals/tasks/ defines one task:
--tag <TAG> lists only the tasks with that tag, and --json prints the list
as JSON. If no task matches, the command stops with “no tasks matched the
requested eval selection”.
If a task file is invalid, points to a dir fixture that doesn’t exist, or
reuses an ID that an earlier file (in file name order) already has,
o4 eval list and o4 eval run skip it and print the
reason as a warning: excluded eval task line. o4 reads only .toml files
directly in evals/tasks/. A missing or invalid evals/corpus.toml is an
error.
Task file reference
All four limits must be greater than zero.
A
git fixture checks out the given commit of the repository in your current
directory into a temporary worktree, so tasks can start from a real past state
of your project.
The check command runs in o4’s most restricted sandbox
tier,
airlock: it can write only inside the task’s folder and has no
network access. A check that needs to download dependencies fails, so use
offline commands such as cargo test --offline. On Linux the check needs
bubblewrap; without it, every task is recorded as skipped with
fixture_error.Run evals
Pick the tasks and one or more models:provider/model references as the rest of o4. See
Choosing a model. o4 needs credentials for each provider,
the same as for a normal session. See
Providers and API keys.
o4 prints a line as each task finishes and a summary at the end:
Choose tasks
With no selection options, o4 runs every task in the corpus. These options narrow it down, and a task must match all of the ones you use:--tag <TAG>: tasks that have this tag. Repeat it to allow several tags; a task needs any one of them.--task <ID>: a task by ID. Repeat it to pick several.--difficulty <LEVEL>:easy,medium, orhard.
Compare models
Repeat--model to run every selected task on each model:
--model references resolve, the command stops with
could not resolve model reference and records nothing.
Limit cost
--budget <USD> stops the run once the total cost of finished tasks reaches
the amount. o4 checks the total after each task, so the last task can take the
run past the budget. A run stopped by the budget ends with status=aborted.
The budget can’t be negative. Each task also has its own limits.cost_cap_usd.
Press Ctrl+C to stop a run. o4 records the task in progress as aborted and
marks the run as aborted.
Other options
--prompt-profile <PROFILE>uses one prompt profile for every task instead of the one o4 picks for each model. The values aremodern-minimal,modern-guided,legacy-guided, andlocal-defensive. See Reasoning and prompt profiles.--enable-mcplets o4 use your configured MCP servers during the run. MCP is off in evals unless you pass it. See MCP servers.
Read the results
o4 eval report shows the latest run as a table, with one row per model, a
column per task, and totals:
--run <ID>shows a different run. The run ID is in the summary line (eval run=1 ...).--baseline <ID>also lists the tasks whose result changed between the baseline run and the selected run. Both runs must use the same corpus version.--jsonprints JSON instead of a table.
pass, fail, or skipped, or aborted when you press
Ctrl+C. The JSON output gives the reason in failure_category, and a
note with details:
o4 eval summary groups every recorded task run by the o4 version that ran
it, and shows how many were attempted and completed (passed), how many
completed with no intervention, and the matching rates. Skipped task runs
don’t count as attempted; aborted ones do. Add --json for JSON output, which
also includes the definition of each measure.
Compare prompt profiles
o4 eval paired runs the selected tasks twice per repeat: a control arm with
the prompt profile o4 picks for each model, and a candidate arm with the
modern-minimal profile. It prints paired eval experiment=<ID> first and
records each arm as its own run. It then compares the two arms on pass rate,
prompt tokens, and other measures, and prints a verdict for each model family:
promote, reject, or insufficient-evidence.
--budget, and --enable-mcp options as
o4 eval run. --budget applies to each arm. --repeats <COUNT> sets how
many control and candidate pairs to run. The default is 3. --json prints
JSON.
With fewer than three complete repeats, the verdict can’t be promote. If a
repeat stops early, for example on a budget or Ctrl+C, the command stops
with an error; the runs it recorded stay in .o4/evals.db. --repeats 0 is
an error.
o4 eval paired runs two arms for every repeat, so with the default 3 repeats
it costs about six times as much as one o4 eval run of the same tasks.