Skip to main content
You can run o4 against a model on your own machine instead of a hosted API. Your code and prompts never leave the machine, and there’s no per-token cost. The tradeoff is that local models are usually smaller and slower than hosted ones, and o4’s agent work (reading files, calling tools, editing code) asks a lot of a model. o4 connects to local models in two ways:
  • Ollama, through the built-in ollama provider.
  • Any local server that speaks the OpenAI Chat Completions API, such as LM Studio, llama.cpp’s llama-server or vLLM, as a custom provider.
o4 doesn’t download or run models itself. Start your runtime first, then point o4 at it.

Ollama

The models listed under Provider: ollama in o4 --list-models are Ollama Cloud models. They run on ollama.com and need an OLLAMA_API_KEY. To use a model on your own Ollama server, add it as shown below.
1

Start Ollama and pull a model

Install Ollama, then pull a model that supports tool calling:
Ollama serves its API on http://localhost:11434 by default.
2

Add the model to o4

Create ~/.o4/models.toml with an entry for the model under the ollama provider:
~/.o4/models.toml
id must match the name Ollama uses for the model, including the tag. base_url must end in /v1, because o4 uses Ollama’s OpenAI-compatible endpoint. Add one [[models]] entry per model.
3

Start o4 with the model

Use the colon form. The model ID contains a colon, so ollama/qwen3-coder:30b doesn’t resolve.
When the base URL points to localhost or 127.0.0.1, o4 sends no API key, so you don’t need OLLAMA_API_KEY for local models. Use one of those two names: in o4 0.2.74, an IPv6 address such as http://[::1]:11434/v1 counts as a remote server and needs a key. Local requests can also run longer before o4 gives up: up to 10 minutes each, compared with 5 minutes for Ollama Cloud. Each entry in models.toml needs id, name, api, provider, base_url and a [models.cost] table. If one entry is missing a field, o4 ignores the whole file, and it also ignores the file if other users can write to it. Run o4 --list-models after each edit and check that your model appears under Provider: ollama. See Choosing a model for more on this file. If your Ollama model supports thinking, add reasoning = true and the levels it accepts so you can set the reasoning effort. Put reasoning with the other model fields, before the first [models.…] table:
~/.o4/models.toml

LM Studio, llama.cpp, vLLM and other local servers

Most local runtimes can serve an OpenAI-compatible API. Add them to ~/.o4/providers.toml as a custom provider. This example uses LM Studio’s default address:
~/.o4/providers.toml
Use your server’s address and port in base_url. The model id must match the name your server reports. o4 requires a key for every custom provider, even when the server doesn’t check it. Set the variable named in api_key_env to any value:
See OpenAI-compatible endpoints for every field, including the compat settings for servers that reject some request fields.
o4 gives each request to a custom provider up to 5 minutes, including the time to stream the reply. A slow local model that takes longer on one reply fails with a timeout. The built-in ollama provider allows 10 minutes for local servers, so prefer it for Ollama.
You can also start this from inside o4: open /config providers and choose Add custom provider. The wizard accepts http://localhost: and http://127.0.0.1: base URLs. It writes the provider to ~/.o4/providers.toml without any models, so you still need to add the [[providers.models]] entries by hand.

Context length

Set context_window to the context length your runtime actually serves, not the model’s advertised maximum. Local runtimes often load models with a much smaller context than the model supports, and o4 has no way to ask the runtime what it chose. o4 uses context_window to show how full the context is and to compact the conversation before it overflows. If the value is larger than the real limit, the runtime can cut off or reject long conversations before o4 compacts them. If you leave context_window out of a custom provider, it defaults to 0 and o4 doesn’t know the limit at all. o4’s system prompt, tool definitions and project instructions take several thousand tokens before your first message, so a model with a small context window fills up fast. To see the size of the system prompt for a model, run o4 -m <model> --print-system-prompt and read the Estimated prompt tokens line at the end. See Context and cost for how compaction works.

How o4 adapts to local models

o4 changes its system prompt for local and unknown models:
  • Models under the ollama provider use the local-defensive prompt profile. Custom providers use legacy-guided.
  • For both, o4 adds a short section of extra guidance to the system prompt. It asks the model to use one tool at a time where it can, keep answers short, search narrowly, and prefer simple plans.
If a strong model on your own server works better with a different profile, override it with --prompt-profile:

Tips for better results

  • Use a model trained for tool calling. o4 does all its work through tools. A model that can’t produce reliable tool calls can chat but can’t read or change your code.
  • Give it as much context as your machine allows. Raise the context length in your runtime’s settings, then set the same value as context_window in o4.
  • Start in plan mode for bigger tasks. A smaller model does better when it plans before editing. See Plan mode.
  • Keep hosted models for hard problems. Switch with /model when a task needs more capability. See Choosing a model.
  • Don’t count on the router to fall back to Ollama. In o4 0.2.74, the model router never switches to an ollama model. A custom provider can be a fallback.