Skip to main content
Many model hosts and local servers accept requests in the format of OpenAI’s Chat Completions API. You can add any of them to o4 as a custom provider: a hosted service such as Together AI, a company gateway, or a server on your own machine. You describe the provider and its models in ~/.o4/providers.toml, and o4 treats them like built-in models. o4 sends each request to <base_url>/chat/completions, the same route OpenAI’s API uses. The endpoint must support that route. Endpoints that only offer OpenAI’s newer Responses API don’t work as custom providers.

Add a provider

1

Describe the provider and its models

Create ~/.o4/providers.toml:
~/.o4/providers.toml
2

Give o4 the key

Export the variable you named in api_key_env:
Or store the key in o4: open /config providers, set Provider to together, and fill in API key. A stored key takes precedence over the environment variable.
3

Check that o4 sees the models

4

Start a session with the model

You can also pick it from /model during a session.
o4 reads providers.toml when it starts, so restart o4 after you edit the file.
If any entry in providers.toml has a TOML error or is missing a required field, o4 skips the whole file. It also skips the file if other users can write to it. o4 records the reason in ~/.o4/o4.log. If your models don’t show up in o4 --list-models, look there first.

Provider fields

Each [[providers]] table defines one provider. You can define as many as you like in the same file. For a service that expects the key in a different header, change both auth fields:
Every custom provider needs a key, either stored in o4 or in the api_key_env variable. If the endpoint doesn’t check keys, set the variable to any placeholder value. Choose an id that isn’t already a built-in provider ID such as openai or anthropic. See Providers and API keys for the built-in list.

Model fields

Each [[providers.models]] table after a provider adds one model to it. Always set context_window to the model’s real limit. o4 uses it to track how full the context is and to decide when to compact the conversation, and the default of 0 means o4 doesn’t know the limit. See Context and cost. Set cost if you want cost totals to reflect this provider:

Adjust requests for strict endpoints

By default o4 sends the usual Chat Completions fields: stream, max_tokens, temperature and tools, when it has values for them. Some endpoints reject fields they don’t support. The compat table on each model controls what o4 sends. The simplest control is an allowlist. When supported_parameters is set, o4 sends only the listed fields:
o4 is a coding agent, and it can only read and edit files through tools. If the endpoint can’t take tool definitions, the model can talk but can’t act on your code. Keep tools in the list unless you’re sure.
Instead of the allowlist, you can turn single fields on or off with booleans:
If you set both, supported_parameters wins. Some servers render chat templates that expect tool-call arguments as a JSON object rather than the JSON string OpenAI uses. They fail with errors such as Can only get item pairs from a mapping. For those, add:
When the endpoint returns 400 or 422, the error in o4 lists the fields o4 sent and your compat settings, and suggests a supported_parameters fix. Use it to find the field to drop.

Reasoning

reasoning = true tells o4 the model produces reasoning. By itself, it doesn’t make o4 send a reasoning effort. To let you choose an effort, declare the levels the endpoint accepts, lowest first:
Valid levels are minimal, low, medium, high, xhigh and max. With supports_reasoning_effort turned on, o4 sends a top-level reasoning_effort field and an extra_body field that turns thinking on. That request shape follows Moonshot’s API, so check that your endpoint accepts it. The field only takes low, high or max: minimal and low go out as low, medium and high as high, and xhigh and max as max. A level the model doesn’t declare first becomes the highest declared level below it. When the reasoning level is default, or the model declares no efforts, o4 sends no reasoning_effort and asks the endpoint to turn thinking off. Without an efforts list, the reasoning setting only offers default.

A complete example

This provider accepts only a few fields, supports reasoning, and needs tool arguments as objects:
~/.o4/providers.toml

Add a provider from the interface

You can also create the provider entry from inside o4. Open /config providers and press a, or choose Add custom provider. The wizard asks for:
  1. Provider ID: 2 to 32 characters, lowercase letters, digits and hyphens, starting with a letter or digit.
  2. Base URL: must start with https://, http://localhost:, or http://127.0.0.1:.
  3. API key environment variable: uppercase letters, digits and underscores, starting with a letter, such as FIREWORKS_API_KEY. The confirm screen shows whether that variable is set.
/config providers add and /add-provider open the Providers tab with the cursor on Add custom provider; press Enter to start the wizard. If the ID is already a custom provider, o4 shows Provider '<id>' already exists and saves nothing. When you save, o4 appends the provider to ~/.o4/providers.toml and shows Provider '<id>' added. Add models by editing ~/.o4/providers.toml. The wizard doesn’t add models, so the provider has none until you add [[providers.models]] entries by hand. The wizard also uses the ID as the display name and the default auth header; edit the file to change them.

Model references with slashes

Many hosted model IDs contain a slash, such as Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8. All of these work:
The bare ID works as long as no other provider has a model with the same ID. Use the provider: prefix to be sure.

How custom providers differ from built-in ones

  • System prompt. o4 treats custom providers as unknown models: it picks the legacy-guided prompt profile and adds extra guidance meant for smaller local models. You can override the profile with --prompt-profile. See Reasoning and prompt profiles.
  • Key privacy. o4 removes the api_key_env variable of every custom provider from the environment of commands it runs for the model, the same as built-in provider keys.
  • Local servers. For LM Studio, llama.cpp, vLLM and similar servers on your machine, see Local models.