Skip to main content
o4 eval runs o4 on a set of coding tasks, checks each result with a command you choose, and records pass or fail, cost, and token use. Use it to compare models, or prompt profiles, on tasks that look like your own work. Evals call real models, so every run costs money. Set a budget before you run a large set.

How it works

An eval corpus is an evals folder in the directory where you run o4 eval. It holds one TOML file per task. For each task and each model you pick, o4:
  1. Copies the task’s fixture (a small project) into a new temporary folder.
  2. Runs o4 non-interactively on the task’s prompt in that folder.
  3. Runs the task’s check command in the folder. The task passes if the command exits with code 0.
  4. Records the result in .o4/evals.db in the directory where you ran the command.
The model works only in the temporary copy, so your real files don’t change.

Set up a corpus

Create this layout in your project:
evals/corpus.toml holds a version string. Change it when you change the tasks in a way that makes old results not comparable:
Each file in evals/tasks/ defines one task:
Check that o4 can read your tasks:
--tag <TAG> lists only the tasks with that tag, and --json prints the list as JSON. If no task matches, the command stops with “no tasks matched the requested eval selection”. If a task file is invalid, points to a dir fixture that doesn’t exist, or reuses an ID that an earlier file (in file name order) already has, o4 eval list and o4 eval run skip it and print the reason as a warning: excluded eval task line. o4 reads only .toml files directly in evals/tasks/. A missing or invalid evals/corpus.toml is an error.

Task file reference

All four limits must be greater than zero. A git fixture checks out the given commit of the repository in your current directory into a temporary worktree, so tasks can start from a real past state of your project.
The check command runs in o4’s most restricted sandbox tier, airlock: it can write only inside the task’s folder and has no network access. A check that needs to download dependencies fails, so use offline commands such as cargo test --offline. On Linux the check needs bubblewrap; without it, every task is recorded as skipped with fixture_error.

Run evals

Pick the tasks and one or more models:
Use the same provider/model references as the rest of o4. See Choosing a model. o4 needs credentials for each provider, the same as for a normal session. See Providers and API keys. o4 prints a line as each task finishes and a summary at the end:

Choose tasks

With no selection options, o4 runs every task in the corpus. These options narrow it down, and a task must match all of the ones you use:
  • --tag <TAG>: tasks that have this tag. Repeat it to allow several tags; a task needs any one of them.
  • --task <ID>: a task by ID. Repeat it to pick several.
  • --difficulty <LEVEL>: easy, medium, or hard.
If nothing matches, the command stops with “no tasks matched the requested eval selection”.

Compare models

Repeat --model to run every selected task on each model:
If a model reference doesn’t resolve, o4 warns and records that model’s tasks as skipped. If a model has no API key, its tasks are also recorded as skipped:
If none of the --model references resolve, the command stops with could not resolve model reference and records nothing.

Limit cost

--budget <USD> stops the run once the total cost of finished tasks reaches the amount. o4 checks the total after each task, so the last task can take the run past the budget. A run stopped by the budget ends with status=aborted. The budget can’t be negative. Each task also has its own limits.cost_cap_usd. Press Ctrl+C to stop a run. o4 records the task in progress as aborted and marks the run as aborted.

Other options

  • --prompt-profile <PROFILE> uses one prompt profile for every task instead of the one o4 picks for each model. The values are modern-minimal, modern-guided, legacy-guided, and local-defensive. See Reasoning and prompt profiles.
  • --enable-mcp lets o4 use your configured MCP servers during the run. MCP is off in evals unless you pass it. See MCP servers.
During an eval, o4 runs without approval prompts, and its shell commands are sandboxed to the task’s temporary folder. Your hooks and permission rules don’t apply.

Read the results

o4 eval report shows the latest run as a table, with one row per model, a column per task, and totals:
  • --run <ID> shows a different run. The run ID is in the summary line (eval run=1 ...).
  • --baseline <ID> also lists the tasks whose result changed between the baseline run and the selected run. Both runs must use the same corpus version.
  • --json prints JSON instead of a table.
A task run ends as pass, fail, or skipped, or aborted when you press Ctrl+C. The JSON output gives the reason in failure_category, and a note with details: o4 eval summary groups every recorded task run by the o4 version that ran it, and shows how many were attempted and completed (passed), how many completed with no intervention, and the matching rates. Skipped task runs don’t count as attempted; aborted ones do. Add --json for JSON output, which also includes the definition of each measure.

Compare prompt profiles

o4 eval paired runs the selected tasks twice per repeat: a control arm with the prompt profile o4 picks for each model, and a candidate arm with the modern-minimal profile. It prints paired eval experiment=<ID> first and records each arm as its own run. It then compares the two arms on pass rate, prompt tokens, and other measures, and prints a verdict for each model family: promote, reject, or insufficient-evidence.
It takes the same selection, --budget, and --enable-mcp options as o4 eval run. --budget applies to each arm. --repeats <COUNT> sets how many control and candidate pairs to run. The default is 3. --json prints JSON. With fewer than three complete repeats, the verdict can’t be promote. If a repeat stops early, for example on a budget or Ctrl+C, the command stops with an error; the runs it recorded stay in .o4/evals.db. --repeats 0 is an error. o4 eval paired runs two arms for every repeat, so with the default 3 repeats it costs about six times as much as one o4 eval run of the same tasks.