> ## Documentation Index
> Fetch the complete documentation index at: https://docs.open4rena.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Model router and fallback

> Switch to another model automatically when a provider is rate limited, overloaded, or failing.

The model router keeps a session working when a provider has trouble. When a request fails with a rate limit, an overload, a timeout, or a server error, o4 switches the session to the next usable model in your list and retries the request there. Without the router, o4 only retries the same model.

The router is off by default. Turn it on when you have keys for more than one provider and don't want a busy provider to stop your work. The router keeps one model per provider, so you need credentials for at least two providers.

## Set up the router

<Steps>
  <Step title="Preview the configuration">
    Run `/router setup` with the models you want to fall back to, in order:

    ```text theme={null}
    /router setup --candidate anthropic:claude-sonnet-4-6 --candidate openai:gpt-5.2
    ```

    Without `--yes`, o4 only shows the router status the configuration would give: the primary model, the usable fallbacks and any skipped models. Without `--candidate`, o4 picks the models itself (see [Which models the router can use](#which-models-the-router-can-use)).

    If fewer than two providers are usable, you see `Router setup was not enabled.` with a hint on which key to add, and o4 writes nothing, even with `--yes`.
  </Step>

  <Step title="Write it">
    Add `--yes` to save it to `~/.o4/config.toml`. Add `--project` to save it to the project's `.o4/config.toml` instead:

    ```text theme={null}
    /router setup --yes --candidate anthropic:claude-sonnet-4-6 --candidate openai:gpt-5.2
    ```

    o4 sets `enabled = true` and replaces the `candidates` list in that file with the models you pass. Without `--candidate`, it removes the list, so o4 goes back to picking the models itself. `--capacity` replaces the `capacity` limits; without it, the limits already in the file stay. Other `[router]` keys stay as they are. o4 doesn't write if a file that takes precedence already has a `[router]` section: the project's `.o4/config.toml` or `.o4/config.local.toml`, or `~/.o4/managed/config.toml`. Edit that file instead.
  </Step>

  <Step title="Start a new session">
    Router changes take effect in the next session.
  </Step>
</Steps>

You can also write the `[router]` section by hand:

```toml theme={null}
[router]
enabled = true
candidates = ["anthropic:claude-sonnet-4-6", "openai:gpt-5.2"]
max_fallback_attempts = 2
```

### From the settings

The **Providers** tab of `/config` has three router rows:

* **Router setup** shows the same preview as `/router setup` with no options. It never writes the configuration; run `/router setup --yes` for that. o4 shows this row while the router is off.
* **Router disable** takes the place of **Router setup** while the router is on. It runs `/router disable`.
* **Router status** runs `/router status`.

## Which models the router can use

The first usable model in the list is the one the session starts on. o4 builds the list like this:

* **With `candidates`:** the model you passed with `-m`, then the session's saved model when you resume, then the models in `candidates`, in order.
* **Without `candidates`:** the model you passed with `-m`, the session's saved model, your default model, the default model of your default provider, then the default model of each other provider you have credentials for.

o4 leaves out any model it can't find in the model catalog, and any model whose provider has no credentials. A built-in `ollama` model that points at a local server (`localhost` or `127.0.0.1`) needs none. The list keeps only the first model of each provider; with `candidates`, `/router status` warns about each extra model from a provider that is already in the list.

With the router on, a model you pass with `-m` that is missing or has no credentials doesn't stop o4. The session starts on the next usable model instead.

<Note>
  Without `candidates`, "each other provider you have credentials for" means providers with a key stored in o4's settings, custom providers whose key variable is set, and the `kimi-coding`, `openai-codex` and `claude-code` providers when o4 finds their credentials. A built-in provider whose key is only in an environment variable, such as `ANTHROPIC_API_KEY`, joins the list only as your default provider. If `/router setup` finds too few providers, store the keys under `/config providers`, or list the models in `candidates`, where environment variables count.
</Note>

When a request fails, o4 also skips a model for this switch if:

* its context window is unknown, or the conversation already fills more than `context_safety_ratio` of it (75% by default);
* it has no tool support, or no reasoning support when the session started with a reasoning level other than `default`. In o4 0.2.74, the router doesn't treat built-in `ollama` models or the Bedrock Kimi models as tool-capable, so it never switches to them;
* it's the model that just failed.

## What happens when a request fails

1. o4 sorts the error into one of four kinds: `rate_limit` (including HTTP 429), `overloaded`, `timeout` (including dropped connections), or `server_error` (HTTP 500, 502 and 503). Any other error, such as a bad API key, never causes a fallback.
2. If that kind is in `fallback_on`, o4 switches to the next usable model in list order and retries the request on it. A notice and the activity row show `Router switched route (<from> -> <to>)` with the reason.
3. If no model is usable, or the kind isn't in `fallback_on`, the activity row shows `Router fallback unavailable` with the reason, and o4 falls back to its normal retries on the current model.
4. After `max_fallback_attempts` switches (1 by default), o4 stops switching and ends the turn with the error. The count starts again after a turn that succeeds.

The router acts only on errors that o4 would retry anyway, and only before the model has streamed any text in that reply. If a reply fails partway through, the turn ends with the error.

After a switch, the session stays on the new model. The router doesn't switch back on its own. To go back, pick the model with `/model`. A model you pick by hand becomes the active model, and the router keeps the rest of the list as fallbacks.

## Limit requests per provider

You can cap how many requests the session sends to one provider at the same time. The cap is shared by all tabs of the interactive interface, and applies to [print mode](/guides/print-mode) runs too. It doesn't cover [subagents](/guides/subagents) or [campaign](/guides/campaigns) workers, and it only works while the router is on:

```toml theme={null}
[router.capacity.anthropic]
max_concurrent_requests = 4
```

Or with `/router setup`:

```text theme={null}
/router setup --yes --capacity anthropic=4
```

Extra requests wait until one finishes.

## Check the router

```text theme={null}
/router status
```

`/router` on its own does the same. It shows whether the router is on, whether a fallback is active, the primary model, the number of providers it can route to, each usable fallback with its context window, each skipped model with the reason (`model not found` or `missing auth`) and a hint on which key to add, any capacity limits, and warnings such as a second model from the same provider. Fallback is active only when the router is on and at least two providers are usable.

`/router status` reads the `[router]` settings the session started with, and your current keys. It shows the list o4 would build now, which can differ from the session's running list, for example after a fallback or a `/model` switch.

## Turn the router off

```text theme={null}
/router disable
```

This sets `enabled = false` in `~/.o4/config.toml` and keeps the rest of the `[router]` section. Add `--project` to change the project's `.o4/config.toml`. `/router off` does the same. The change takes effect in the next session. If a file that takes precedence sets `enabled = true`, the router stays on; set it there instead.

## Settings

All `[router]` settings, with their types and defaults, are in the [configuration reference](/reference/configuration#router):

| Setting | Default | What it does |
| - | - | - |
| `enabled` | `false` | Turns the router on. |
| `mode` | `session` | Only `session` is supported. Any other value turns the router off, with a warning in `~/.o4/o4.log`. |
| `candidates` | none | The models to fall back to, in order. Without it, o4 builds the list itself. |
| `max_fallback_attempts` | `1` | How many times o4 may switch before the turn fails. |
| `context_safety_ratio` | `0.75` | How much of a model's context window the conversation may fill for it to stay a candidate. |
| `fallback_on` | all four kinds | Which errors cause a fallback: `rate_limit`, `overloaded`, `timeout`, `server_error`. |
| `capacity.<provider>.max_concurrent_requests` | none | A cap on requests in flight to one provider. `<provider>` is a provider ID such as `anthropic`. |

o4 checks these values when a session starts. It replaces an out-of-range `max_fallback_attempts` or `context_safety_ratio` with the default, drops unknown `fallback_on` entries and capacity limits below 1, and logs a warning for each.
