Skip to main content
Every model has a context window: a limit on how much text it can take in one request. In o4 that text is your conversation, the files and command output the agent has read, the system prompt, and the tool definitions. As a session grows, it fills the window and each request costs more. This page shows how to see where you stand and how o4 makes room.

Watch the context meter

While you work, the prompt box shows ctx and a percentage, for example ctx 42%. It’s the share of the model’s context window that the next request will use. It turns yellow at 75% and red at 90%. When the context reaches 80%, o4 also shows a one-time warning: Context at 80% — will auto-compact before next message.

See what fills the context

/context opens a pane that breaks down the context window:
  • A bar for the whole window, with the tokens used out of the total.
  • A bar for each part of the request: instructions (the system prompt and project instructions), messages, tools (tool definitions), tool results, images, current draft, and output reserve (room kept for the reply), plus free space.
  • The numbers behind the bars, including the compaction threshold and whether you’re within capacity.
The pane has three tabs, Context, Session and Optimize. Switch between them with Tab or the arrow keys. Optimize suggests what to do next, for example running /compact when you’re past the compaction threshold, /scout when tool definitions take up a large share of the context, or /reload when your project instructions have changed. Press r to refresh the numbers and Esc or q to close the pane. The numbers come from the last request o4 prepared, so before your first message there’s nothing to show. Send a message, then press r.

See what a session costs

/cost shows the token usage and cost of the current session: o4 works out the cost from the token counts the provider reports and the per-million-token prices in its model catalog. It’s an estimate: your provider’s bill is the real number. With a subscription or a local model, the prices in the catalog may not reflect what you pay. The counts start at zero each time you start or resume o4. When you quit, o4 prints the session’s totals along with the command to resume it. See Sessions. Two smaller commands aren’t in the /help list: /usage prints the same totals on one line, and /tokens <text> estimates how many tokens some text uses, at about 4 characters per token. See Slash commands. o4 has no spending limit for a session: it doesn’t stop when the cost reaches an amount. The only token budget is the optional one on a goal: when it’s used up, the goal becomes budget limited and o4 stops continuing it on its own. /status shows the current model and its context window. For ChatGPT subscriptions it also shows your account’s usage limits; other providers don’t report them. See Claude and ChatGPT subscriptions.

Usage across sessions

/stats opens the same pane as /context, on the Session tab. It shows statistics that o4 records on your machine in ~/.o4/telemetry/events.jsonl: sessions, tool calls and their success rates, retries and errors. Press r to rebuild the tab from the full report. This file is only written while telemetry is on. If you’ve turned telemetry off, /stats has little or nothing to show. See Telemetry and privacy.

Compact the conversation

Compaction replaces older parts of the conversation with a summary, so the model keeps the thread of the work while using far less context. o4 asks the current model to write the summary. After compaction, the model’s context holds:
  • The system prompt and project instructions, unchanged.
  • Your most recent messages, word for word (up to 5, within a size limit).
  • A summary of everything else: the task, the files that were edited, decisions, and errors. Long tool output is trimmed before it’s summarized.
The transcript on your screen doesn’t change. Only what the model sees is replaced. To keep a full record, use /export before compacting. See Sessions.

Compact by hand

Run /compact at any time, for example when you finish one task and start another in the same session:
o4 reports the result in the transcript, for example Context compacted: 48 -> 7 messages. If there’s too little to compact, or the summary wouldn’t save enough, the conversation is left as it is and o4 tells you why. /compact doesn’t take arguments, so you can’t tell it what to focus on. To steer what’s kept, ask the model to write down what matters first, or start a new session with /clear and a short description of where you are.

Automatic compaction

o4 compacts on its own when the context passes the compaction threshold, 70% of the window by default. It checks before sending each message and again after each turn, and shows Summarized earlier context: 52 -> 8 messages. when it compacts. The agent loop also has two safety nets that work even when automatic compaction is off, and in print mode:
  • If the request is about to reach 85% of the window during a turn, it’s compacted before it’s sent.
  • If the provider rejects a request because the context is too long, o4 compacts and tries again.

Change the thresholds

The thresholds are settings in ~/.o4/settings.json:
A project can set them in .o4/settings.json. Project values are used only in a trusted workspace. o4 reads these keys only from settings.json; in config.toml they have no effect. Changes take effect the next time you start o4.
A lower threshold keeps requests smaller and cheaper, but summarizes sooner. A higher one keeps more detail, at the cost of larger requests. If you often work with very large files, a lower threshold leaves more room for them.

How tokens are counted

When the provider reports token counts, o4 uses them. That covers the cost figures in /cost. Before a request is sent, o4 has to estimate its size: it counts about one token for every four characters of text, adds an estimate for each image based on its size, and adds the room reserved for the model’s reply. The context meter and the compaction thresholds use this estimate. It’s close enough to decide when to compact, but it can differ from the provider’s count, especially for code and non-English text. In the /context pane and its JSON output, each number is marked with where it came from, such as provider-reported, estimated, or configured.

Keep context small

  • Start a new session for unrelated work instead of adding to a long one.
  • Point the agent at specific files and functions rather than asking it to explore the whole project.
  • Use /hey for side questions. The side chat doesn’t add to the main context. See Writing prompts.
  • Run /compact between tasks.
  • If /context shows tool definitions taking a large share, run /scout, which suggests a smaller set of tools for the project. See Code intelligence.

Diagnostics from the command line

/context and /stats also work in print mode, and can print JSON for scripts:
Inside the interface, --json isn’t available and shows an error. See Print mode and scripting.