> ## Documentation Index
> Fetch the complete documentation index at: https://docs.open4rena.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI-compatible endpoints

> Connect o4 to any endpoint that speaks the OpenAI API.

Many model hosts and local servers accept requests in the format of OpenAI's Chat Completions API. You can add any of them to o4 as a custom provider: a hosted service such as Together AI, a company gateway, or a server on your own machine. You describe the provider and its models in `~/.o4/providers.toml`, and o4 treats them like built-in models.

o4 sends each request to `<base_url>/chat/completions`, the same route OpenAI's API uses. The endpoint must support that route. Endpoints that only offer OpenAI's newer Responses API don't work as custom providers.

## Add a provider

<Steps>
  <Step title="Describe the provider and its models">
    Create `~/.o4/providers.toml`:

    ```toml ~/.o4/providers.toml theme={null}
    [[providers]]
    id = "together"
    name = "Together AI"
    api_key_env = "TOGETHER_API_KEY"
    base_url = "https://api.together.xyz/v1"

    [[providers.models]]
    id = "Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8"
    name = "Qwen3 Coder 480B"
    context_window = 262144
    max_tokens = 32768
    ```
  </Step>

  <Step title="Give o4 the key">
    Export the variable you named in `api_key_env`:

    ```bash theme={null}
    export TOGETHER_API_KEY="your-key"
    ```

    Or store the key in o4: open `/config providers`, set **Provider** to `together`, and fill in **API key**. A stored key takes precedence over the environment variable.
  </Step>

  <Step title="Check that o4 sees the models">
    ```bash theme={null}
    o4 --list-models
    ```

    ```text theme={null}
    Provider: together
      Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 (context: 262k, max_tokens: 32768, input: text, reasoning: no)
    ```
  </Step>

  <Step title="Start a session with the model">
    ```bash theme={null}
    o4 -m together:Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
    ```

    You can also pick it from `/model` during a session.
  </Step>
</Steps>

o4 reads `providers.toml` when it starts, so restart o4 after you edit the file.

<Warning>
  If any entry in `providers.toml` has a TOML error or is missing a required field, o4 skips the whole file. It also skips the file if other users can write to it. o4 records the reason in `~/.o4/o4.log`. If your models don't show up in `o4 --list-models`, look there first.
</Warning>

## Provider fields

Each `[[providers]]` table defines one provider. You can define as many as you like in the same file.

| Field | Required | Description |
| - | - | - |
| `id` | Yes | Provider ID. Use it in model references such as `id:model`. Don't use `groq`, `mistral`, `openrouter` or `cohere`: o4 removes providers with those IDs, so their models never appear. |
| `name` | Yes | Display name, used in error messages. |
| `api_key_env` | Yes | Name of the environment variable that holds the key. |
| `base_url` | Yes | API root, without `/chat/completions`. For most services this ends in `/v1`. |
| `auth_header` | No | Header that carries the key. Default: `Authorization`. |
| `auth_format` | No | Header value. `{key}` is replaced with the key. Default: `Bearer {key}`. |

For a service that expects the key in a different header, change both auth fields:

```toml theme={null}
auth_header = "api-key"
auth_format = "{key}"
```

Every custom provider needs a key, either stored in o4 or in the `api_key_env` variable. If the endpoint doesn't check keys, set the variable to any placeholder value.

Choose an `id` that isn't already a built-in provider ID such as `openai` or `anthropic`. See [Providers and API keys](/models/providers) for the built-in list.

## Model fields

Each `[[providers.models]]` table after a provider adds one model to it.

| Field | Required | Default | Description |
| - | - | - | - |
| `id` | Yes | | Model name the endpoint expects. o4 sends it unchanged. |
| `name` | Yes | | Display name. |
| `context_window` | No | `0` | Context window in tokens. |
| `max_tokens` | No | `0` | Maximum output tokens. `0` leaves it to the endpoint. |
| `input` | No | `["text"]` | Input types: `"text"`, and `"image"` if the model accepts images. |
| `reasoning` | No | `false` | Whether the model produces reasoning. See [Reasoning](#reasoning). |
| `cost` | No | all `0` | Prices in US dollars per million tokens: `input`, `output`, `cache_read`, `cache_write`. |
| `compat` | No | | Request adjustments for endpoints that reject some fields. See below. |

Always set `context_window` to the model's real limit. o4 uses it to track how full the context is and to decide when to compact the conversation, and the default of `0` means o4 doesn't know the limit. See [Context and cost](/guides/context-and-cost).

Set `cost` if you want [cost totals](/guides/context-and-cost) to reflect this provider:

```toml theme={null}
[providers.models.cost]
input = 0.30
output = 1.20
```

## Adjust requests for strict endpoints

By default o4 sends the usual Chat Completions fields: `stream`, `max_tokens`, `temperature` and `tools`, when it has values for them. Some endpoints reject fields they don't support. The `compat` table on each model controls what o4 sends.

The simplest control is an allowlist. When `supported_parameters` is set, o4 sends only the listed fields:

```toml theme={null}
[providers.models.compat]
supported_parameters = ["stream", "max_tokens", "tools"]
```

| Value | Effect |
| - | - |
| `stream` or `streaming` | Send `stream: true`. |
| `max_tokens` | Send `max_tokens`. |
| `temperature` | Send `temperature`. |
| `tools`, `tool_choice` or `function_call` | Send the tool definitions. |
| `reasoning_content` or `reasoning` | Send the model's earlier reasoning back in the conversation. |
| `reasoning_effort` | Send a reasoning effort. See [Reasoning](#reasoning). |

<Warning>
  o4 is a coding agent, and it can only read and edit files through tools. If the endpoint can't take tool definitions, the model can talk but can't act on your code. Keep `tools` in the list unless you're sure.
</Warning>

Instead of the allowlist, you can turn single fields on or off with booleans:

```toml theme={null}
[providers.models.compat]
supports_streaming = true
supports_max_tokens = true
supports_temperature = false
supports_tools = true
supports_reasoning_content = false
supports_reasoning_effort = false
```

If you set both, `supported_parameters` wins.

Some servers render chat templates that expect tool-call arguments as a JSON object rather than the JSON string OpenAI uses. They fail with errors such as `Can only get item pairs from a mapping`. For those, add:

```toml theme={null}
[providers.models.compat]
tool_call_arguments_object = true
```

When the endpoint returns `400` or `422`, the error in o4 lists the fields o4 sent and your `compat` settings, and suggests a `supported_parameters` fix. Use it to find the field to drop.

## Reasoning

`reasoning = true` tells o4 the model produces reasoning. By itself, it doesn't make o4 send a reasoning effort.

To let you [choose an effort](/models/reasoning), declare the levels the endpoint accepts, lowest first:

```toml theme={null}
[providers.models.compat]
supports_reasoning_effort = true

[providers.models.compat.reasoning]
efforts = ["low", "high"]
```

Valid levels are `minimal`, `low`, `medium`, `high`, `xhigh` and `max`. With `supports_reasoning_effort` turned on, o4 sends a top-level `reasoning_effort` field and an `extra_body` field that turns thinking on. That request shape follows Moonshot's API, so check that your endpoint accepts it. The field only takes `low`, `high` or `max`: `minimal` and `low` go out as `low`, `medium` and `high` as `high`, and `xhigh` and `max` as `max`. A level the model doesn't declare first becomes the highest declared level below it.

When the reasoning level is `default`, or the model declares no `efforts`, o4 sends no `reasoning_effort` and asks the endpoint to turn thinking off. Without an `efforts` list, the reasoning setting only offers `default`.

## A complete example

This provider accepts only a few fields, supports reasoning, and needs tool arguments as objects:

```toml ~/.o4/providers.toml theme={null}
[[providers]]
id = "xiaomi"
name = "Xiaomi MiMo Token Plan"
api_key_env = "MIMO_API_KEY"
base_url = "https://token-plan-sgp.xiaomimimo.com/v1"

[[providers.models]]
id = "MiMo-V2.5-Pro"
name = "MiMo V2.5 Pro"
reasoning = true
input = ["text"]
context_window = 258000
max_tokens = 8192

[providers.models.compat]
supported_parameters = ["stream", "max_tokens", "tools"]
supports_reasoning_content = false
tool_call_arguments_object = true
```

```bash theme={null}
export MIMO_API_KEY="your-token-plan-key"
o4 -m xiaomi/MiMo-V2.5-Pro
```

## Add a provider from the interface

You can also create the provider entry from inside o4. Open `/config providers` and press `a`, or choose **Add custom provider**. The wizard asks for:

1. **Provider ID**: 2 to 32 characters, lowercase letters, digits and hyphens, starting with a letter or digit.
2. **Base URL**: must start with `https://`, `http://localhost:`, or `http://127.0.0.1:`.
3. **API key environment variable**: uppercase letters, digits and underscores, starting with a letter, such as `FIREWORKS_API_KEY`. The confirm screen shows whether that variable is set.

`/config providers add` and `/add-provider` open the **Providers** tab with the cursor on **Add custom provider**; press `Enter` to start the wizard. If the ID is already a custom provider, o4 shows `Provider '<id>' already exists` and saves nothing.

When you save, o4 appends the provider to `~/.o4/providers.toml` and shows `Provider '<id>' added. Add models by editing ~/.o4/providers.toml.`

The wizard doesn't add models, so the provider has none until you add `[[providers.models]]` entries by hand. The wizard also uses the ID as the display name and the default auth header; edit the file to change them.

## Model references with slashes

Many hosted model IDs contain a slash, such as `Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8`. All of these work:

```bash theme={null}
o4 -m together:Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
o4 -m together/Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
o4 -m Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
```

The bare ID works as long as no other provider has a model with the same ID. Use the `provider:` prefix to be sure.

## How custom providers differ from built-in ones

* **System prompt.** o4 treats custom providers as unknown models: it picks the `legacy-guided` prompt profile and adds extra guidance meant for smaller local models. You can override the profile with `--prompt-profile`. See [Reasoning and prompt profiles](/models/reasoning#prompt-profiles).
* **Key privacy.** o4 removes the `api_key_env` variable of every custom provider from the environment of commands it runs for the model, the same as built-in provider keys.
* **Local servers.** For LM Studio, llama.cpp, vLLM and similar servers on your machine, see [Local models](/models/local-models).

## Related

* [Providers and API keys](/models/providers)
* [Choosing a model](/models/overview)
* [Configuration reference](/reference/configuration)
