> ## Documentation Index
> Fetch the complete documentation index at: https://docs.open4rena.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Local models

> Run o4 against models on your own machine.

You can run o4 against a model on your own machine instead of a hosted API. Your code and prompts never leave the machine, and there's no per-token cost. The tradeoff is that local models are usually smaller and slower than hosted ones, and o4's agent work (reading files, calling tools, editing code) asks a lot of a model.

o4 connects to local models in two ways:

* **Ollama**, through the built-in `ollama` provider.
* **Any local server that speaks the OpenAI Chat Completions API**, such as LM Studio, llama.cpp's `llama-server` or vLLM, as a custom provider.

o4 doesn't download or run models itself. Start your runtime first, then point o4 at it.

## Ollama

<Note>
  The models listed under `Provider: ollama` in `o4 --list-models` are Ollama Cloud models. They run on `ollama.com` and need an `OLLAMA_API_KEY`. To use a model on your own Ollama server, add it as shown below.
</Note>

<Steps>
  <Step title="Start Ollama and pull a model">
    Install [Ollama](https://ollama.com), then pull a model that supports tool calling:

    ```bash theme={null}
    ollama pull qwen3-coder:30b
    ```

    Ollama serves its API on `http://localhost:11434` by default.
  </Step>

  <Step title="Add the model to o4">
    Create `~/.o4/models.toml` with an entry for the model under the `ollama` provider:

    ```toml ~/.o4/models.toml theme={null}
    [[models]]
    id = "qwen3-coder:30b"
    name = "Qwen3 Coder 30B (local)"
    api = "ollama-native"
    provider = "ollama"
    base_url = "http://localhost:11434/v1"
    input = ["text"]
    context_window = 32768
    max_tokens = 8192

    [models.cost]
    input = 0.0
    output = 0.0
    ```

    `id` must match the name Ollama uses for the model, including the tag. `base_url` must end in `/v1`, because o4 uses Ollama's OpenAI-compatible endpoint. Add one `[[models]]` entry per model.
  </Step>

  <Step title="Start o4 with the model">
    ```bash theme={null}
    o4 -m ollama:qwen3-coder:30b
    ```

    Use the colon form. The model ID contains a colon, so `ollama/qwen3-coder:30b` doesn't resolve.
  </Step>
</Steps>

When the base URL points to `localhost` or `127.0.0.1`, o4 sends no API key, so you don't need `OLLAMA_API_KEY` for local models. Use one of those two names: in o4 0.2.74, an IPv6 address such as `http://[::1]:11434/v1` counts as a remote server and needs a key. Local requests can also run longer before o4 gives up: up to 10 minutes each, compared with 5 minutes for Ollama Cloud.

Each entry in `models.toml` needs `id`, `name`, `api`, `provider`, `base_url` and a `[models.cost]` table. If one entry is missing a field, o4 ignores the whole file, and it also ignores the file if other users can write to it. Run `o4 --list-models` after each edit and check that your model appears under `Provider: ollama`. See [Choosing a model](/models/overview#add-or-change-model-entries) for more on this file.

If your Ollama model supports thinking, add `reasoning = true` and the levels it accepts so you can [set the reasoning effort](/models/reasoning). Put `reasoning` with the other model fields, before the first `[models.…]` table:

```toml ~/.o4/models.toml theme={null}
[[models]]
id = "qwen3-coder:30b"
name = "Qwen3 Coder 30B (local)"
api = "ollama-native"
provider = "ollama"
base_url = "http://localhost:11434/v1"
input = ["text"]
context_window = 32768
max_tokens = 8192
reasoning = true

[models.cost]
input = 0.0
output = 0.0

[models.compat.reasoning]
efforts = ["low", "medium", "high"]
```

## LM Studio, llama.cpp, vLLM and other local servers

Most local runtimes can serve an OpenAI-compatible API. Add them to `~/.o4/providers.toml` as a custom provider. This example uses LM Studio's default address:

```toml ~/.o4/providers.toml theme={null}
[[providers]]
id = "lmstudio"
name = "LM Studio"
api_key_env = "LMSTUDIO_API_KEY"
base_url = "http://localhost:1234/v1"

[[providers.models]]
id = "qwen2.5-coder-7b-instruct"
name = "Qwen2.5 Coder 7B"
context_window = 32768
max_tokens = 8192
```

Use your server's address and port in `base_url`. The model `id` must match the name your server reports.

o4 requires a key for every custom provider, even when the server doesn't check it. Set the variable named in `api_key_env` to any value:

```bash theme={null}
export LMSTUDIO_API_KEY=local
o4 -m lmstudio:qwen2.5-coder-7b-instruct
```

See [OpenAI-compatible endpoints](/models/openai-compatible) for every field, including the `compat` settings for servers that reject some request fields.

<Note>
  o4 gives each request to a custom provider up to 5 minutes, including the time to stream the reply. A slow local model that takes longer on one reply fails with a timeout. The built-in `ollama` provider allows 10 minutes for local servers, so prefer it for Ollama.
</Note>

You can also start this from inside o4: open `/config providers` and choose **Add custom provider**. The wizard accepts `http://localhost:` and `http://127.0.0.1:` base URLs. It writes the provider to `~/.o4/providers.toml` without any models, so you still need to add the `[[providers.models]]` entries by hand.

## Context length

Set `context_window` to the context length your runtime actually serves, not the model's advertised maximum. Local runtimes often load models with a much smaller context than the model supports, and o4 has no way to ask the runtime what it chose.

o4 uses `context_window` to show how full the context is and to compact the conversation before it overflows. If the value is larger than the real limit, the runtime can cut off or reject long conversations before o4 compacts them. If you leave `context_window` out of a custom provider, it defaults to `0` and o4 doesn't know the limit at all.

o4's system prompt, tool definitions and project instructions take several thousand tokens before your first message, so a model with a small context window fills up fast. To see the size of the system prompt for a model, run `o4 -m <model> --print-system-prompt` and read the `Estimated prompt tokens` line at the end. See [Context and cost](/guides/context-and-cost) for how compaction works.

## How o4 adapts to local models

o4 changes its system prompt for local and unknown models:

* Models under the `ollama` provider use the `local-defensive` [prompt profile](/models/reasoning#prompt-profiles). Custom providers use `legacy-guided`.
* For both, o4 adds a short section of extra guidance to the system prompt. It asks the model to use one tool at a time where it can, keep answers short, search narrowly, and prefer simple plans.

If a strong model on your own server works better with a different profile, override it with `--prompt-profile`:

```bash theme={null}
o4 -m lmstudio:qwen2.5-coder-7b-instruct --prompt-profile modern-guided
```

## Tips for better results

* **Use a model trained for tool calling.** o4 does all its work through tools. A model that can't produce reliable tool calls can chat but can't read or change your code.
* **Give it as much context as your machine allows.** Raise the context length in your runtime's settings, then set the same value as `context_window` in o4.
* **Start in plan mode for bigger tasks.** A smaller model does better when it plans before editing. See [Plan mode](/guides/plan-mode).
* **Keep hosted models for hard problems.** Switch with `/model` when a task needs more capability. See [Choosing a model](/models/overview).
* **Don't count on the router to fall back to Ollama.** In o4 0.2.74, the [model router](/models/router) never switches to an `ollama` model. A custom provider can be a fallback.

## Related

* [OpenAI-compatible endpoints](/models/openai-compatible)
* [Providers and API keys](/models/providers)
* [Reasoning and prompt profiles](/models/reasoning)
