Skip to main content
A harness file tells o4 how work should be done in a project. It lives at .o4/harness.toml in the directory you start o4 in (o4 doesn’t look in parent folders) and can hold:
  • Guides: instructions o4 adds to the system prompt, such as coding rules or a review checklist.
  • Sensors: commands that check the work, such as tests or a linter.
  • Workflows: which guides and sensors apply to which kind of task.
  • Completion policies: what the model should show before it says the work is done.
  • Escalation rules: accepted, but not used in 0.2.74.
For standing instructions that apply to every task, project instructions in AGENTS.md are simpler. Use a harness when different kinds of tasks need different guides, or when campaigns should check their results with commands.
o4 reads .o4/harness.toml only in a trusted workspace, because guides enter the system prompt and sensors are commands. In an untrusted workspace, o4 ignores the file without a warning. The file can be up to 1 MiB.

A first harness

In an interactive session, o4 uses the workflow named in [defaults]. It adds that workflow’s guides to the system prompt, lists its sensors with their commands so the model knows how to check its work, and lists what its completion policy requires. Without a default workflow, o4 uses every guide and sensor in the file. o4 reads the file at startup, so restart o4 after you change it. Don’t use /reload for this: it rebuilds the system prompt without the harness section, so the guides drop out. To see what the model gets, run o4 --print-system-prompt and look for the Project Harness Guides section.

How o4 uses the harness

o4 doesn’t run sensors on its own in an interactive session or in print mode. The model sees the commands and can run them with its bash tool, under the same permissions as any other command. In a campaign, o4 runs the sensors after the workers finish (with the sequential strategy, also after each worker), in the project directory, without asking, and doesn’t retry them. A blocking sensor that fails, times out or can’t start fails every worker’s validation, and the campaign ends as failed. A non-blocking sensor’s result is recorded but doesn’t fail anything. The results, with up to 40 lines of each command’s output and its fix_hint, are saved in the campaign’s result.json under .o4/campaign/. Campaigns don’t check completion policies or escalation rules in 0.2.74.

Task types

Print mode and campaigns pick a task type from the prompt or the campaign objective. o4 checks the keywords below in order and uses the first row that matches. Matching ignores case and also matches inside longer words, so “latest” counts as test: o4 uses the first workflow whose task_types includes the detected type. If none matches, it uses the default workflow. If there’s no default workflow either, it uses every guide and sensor in the file.

Reference

defaults

[defaults] is optional.

guides

Each [[guides]] entry is one guide. If an optional guide’s file is missing or can’t be read, or the guide has neither path nor text, o4 skips the guide and adds a warning to the system prompt. If a required guide has one of those problems, o4 skips the whole harness and logs the error to ~/.o4/o4.log. o4 also skips a guide whose text already appears in your project instructions, so it isn’t sent twice.
If .o4/harness.toml isn’t valid TOML, or a required key is missing, o4 skips the whole harness and doesn’t log anything. Run the harness_audit tool or check for the Project Harness Guides section in o4 --print-system-prompt to catch this.

sensors

Each [[sensors]] entry is one check command. An empty name or command, two sensors with the same name, or an invalid paths pattern makes o4 skip the whole harness and log the error to ~/.o4/o4.log.

workflows

completion_policies

A completion policy lists what the model should show before it says the work is done. When the selected workflow names a policy, o4 adds these requirements to the system prompt, one line per key set to true. It’s guidance for the model: o4 doesn’t check it. All the require_ keys are false by default.

escalation_rules

Accepted, but not used in 0.2.74: o4 reads [[escalation_rules]] entries and doesn’t act on them. Each entry needs a name and a condition, or o4 skips the whole harness. The conditions the code knows are blocking_sensor_failed, missing_required_evidence, snapshot_or_golden_changed, unresolved_risks and retry_limit_exhausted. The other keys are action (ask_user by default) and message.

Check a harness file

o4 has no command for this. Ask the model to audit the harness instead. Its harness_audit tool reads .o4/harness.toml without changing it and never asks for approval. See Tools.
The tool reports each finding as an error, a warning or a suggestion: If the file doesn’t parse, the tool returns the parse error. In an untrusted workspace it reports that no harness file was found. The tool has two optional inputs, project_path to audit another folder and history with recent guide and sensor results, which adds findings for guides and sensors that didn’t appear and sensors that failed at least twice.