> ## Documentation Index
> Fetch the complete documentation index at: https://docs.open4rena.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Context and cost

> See how much of the context window and budget a session uses, and compact it.

Every model has a context window: a limit on how much text it can take in one request. In o4 that text is your conversation, the files and command output the agent has read, the system prompt, and the tool definitions. As a session grows, it fills the window and each request costs more. This page shows how to see where you stand and how o4 makes room.

## Watch the context meter

While you work, the prompt box shows `ctx` and a percentage, for example `ctx 42%`. It's the share of the model's context window that the next request will use. It turns yellow at 75% and red at 90%.

When the context reaches 80%, o4 also shows a one-time warning: `Context at 80% — will auto-compact before next message`.

## See what fills the context

`/context` opens a pane that breaks down the context window:

* A bar for the whole window, with the tokens used out of the total.
* A bar for each part of the request: `instructions` (the system prompt and project instructions), `messages`, `tools` (tool definitions), `tool results`, `images`, `current draft`, and `output reserve` (room kept for the reply), plus `free space`.
* The numbers behind the bars, including the compaction threshold and whether you're within capacity.

The pane has three tabs, `Context`, `Session` and `Optimize`. Switch between them with `Tab` or the arrow keys. `Optimize` suggests what to do next, for example running `/compact` when you're past the compaction threshold, `/scout` when tool definitions take up a large share of the context, or `/reload` when your project instructions have changed. Press `r` to refresh the numbers and `Esc` or `q` to close the pane.

The numbers come from the last request o4 prepared, so before your first message there's nothing to show. Send a message, then press `r`.

## See what a session costs

`/cost` shows the token usage and cost of the current session:

| Row | Meaning |
| - | - |
| `Input tokens` | Tokens sent to the model. |
| `Output tokens` | Tokens the model wrote. |
| `Cache read tokens` | Input served from the provider's prompt cache, which is usually cheaper. |
| `Cache write tokens` | Input written to the prompt cache. |
| `Total cost` | The estimated cost in US dollars. |
| `Side chat cost` | The part of the cost from `/hey`, when you've used it. |

o4 works out the cost from the token counts the provider reports and the per-million-token prices in its model catalog. It's an estimate: your provider's bill is the real number. With a subscription or a local model, the prices in the catalog may not reflect what you pay. The counts start at zero each time you start or resume o4.

When you quit, o4 prints the session's totals along with the command to resume it. See [Sessions](/guides/sessions).

Two smaller commands aren't in the `/help` list: `/usage` prints the same totals on one line, and `/tokens <text>` estimates how many tokens some text uses, at about 4 characters per token. See [Slash commands](/reference/slash-commands#not-listed-in-/help).

o4 has no spending limit for a session: it doesn't stop when the cost reaches an amount. The only token budget is the optional one on a [goal](/guides/goals#token-budgets): when it's used up, the goal becomes `budget limited` and o4 stops continuing it on its own.

`/status` shows the current model and its context window. For ChatGPT subscriptions it also shows your account's usage limits; other providers don't report them. See [Claude and ChatGPT subscriptions](/models/subscriptions).

## Usage across sessions

`/stats` opens the same pane as `/context`, on the `Session` tab. It shows statistics that o4 records on your machine in `~/.o4/telemetry/events.jsonl`: sessions, tool calls and their success rates, retries and errors. Press `r` to rebuild the tab from the full report.

This file is only written while telemetry is on. If you've turned telemetry off, `/stats` has little or nothing to show. See [Telemetry and privacy](/help/telemetry-and-privacy).

## Compact the conversation

Compaction replaces older parts of the conversation with a summary, so the model keeps the thread of the work while using far less context. o4 asks the current model to write the summary.

After compaction, the model's context holds:

* The system prompt and project instructions, unchanged.
* Your most recent messages, word for word (up to 5, within a size limit).
* A summary of everything else: the task, the files that were edited, decisions, and errors. Long tool output is trimmed before it's summarized.

The transcript on your screen doesn't change. Only what the model sees is replaced. To keep a full record, use `/export` before compacting. See [Sessions](/guides/sessions).

### Compact by hand

Run `/compact` at any time, for example when you finish one task and start another in the same session:

```text theme={null}
/compact
```

o4 reports the result in the transcript, for example `Context compacted: 48 -> 7 messages.` If there's too little to compact, or the summary wouldn't save enough, the conversation is left as it is and o4 tells you why.

`/compact` doesn't take arguments, so you can't tell it what to focus on. To steer what's kept, ask the model to write down what matters first, or start a new session with `/clear` and a short description of where you are.

### Automatic compaction

o4 compacts on its own when the context passes the compaction threshold, 70% of the window by default. It checks before sending each message and again after each turn, and shows `Summarized earlier context: 52 -> 8 messages.` when it compacts.

The agent loop also has two safety nets that work even when automatic compaction is off, and in [print mode](/guides/print-mode):

* If the request is about to reach 85% of the window during a turn, it's compacted before it's sent.
* If the provider rejects a request because the context is too long, o4 compacts and tries again.

### Change the thresholds

The thresholds are settings in `~/.o4/settings.json`:

```json theme={null}
{
  "auto_compact_enabled": true,
  "auto_compact_threshold": 70,
  "auto_compact_warning_threshold": 80
}
```

| Setting | Default | Meaning |
| - | - | - |
| `auto_compact_enabled` | `true` | Compact automatically. You can also switch this with `Auto-compact context` in the setup wizard (`/onboard`). |
| `auto_compact_threshold` | `70` | Percent of the context window at which o4 compacts. |
| `auto_compact_warning_threshold` | `80` | Percent at which o4 shows the one-time warning. |

A project can set them in `.o4/settings.json`. Project values are used only in a [trusted workspace](/safety/workspace-trust). o4 reads these keys only from `settings.json`; in `config.toml` they have no effect. Changes take effect the next time you start o4.

<Tip>
  A lower threshold keeps requests smaller and cheaper, but summarizes sooner. A higher one keeps more detail, at the cost of larger requests. If you often work with very large files, a lower threshold leaves more room for them.
</Tip>

## How tokens are counted

When the provider reports token counts, o4 uses them. That covers the cost figures in `/cost`. Before a request is sent, o4 has to estimate its size: it counts about one token for every four characters of text, adds an estimate for each image based on its size, and adds the room reserved for the model's reply. The context meter and the compaction thresholds use this estimate. It's close enough to decide when to compact, but it can differ from the provider's count, especially for code and non-English text.

In the `/context` pane and its JSON output, each number is marked with where it came from, such as provider-reported, estimated, or configured.

## Keep context small

* Start a new session for unrelated work instead of adding to a long one.
* Point the agent at specific files and functions rather than asking it to explore the whole project.
* Use `/hey` for side questions. The side chat doesn't add to the main context. See [Writing prompts](/guides/prompting).
* Run `/compact` between tasks.
* If `/context` shows tool definitions taking a large share, run `/scout`, which suggests a smaller set of tools for the project. See [Code intelligence](/guides/code-intelligence).

## Diagnostics from the command line

`/context` and `/stats` also work in print mode, and can print JSON for scripts:

```bash theme={null}
o4 -p "/context --json"
o4 -p "/stats --json"
```

Inside the interface, `--json` isn't available and shows an error. See [Print mode and scripting](/guides/print-mode).
