Skip to main content
The model router keeps a session working when a provider has trouble. When a request fails with a rate limit, an overload, a timeout, or a server error, o4 switches the session to the next usable model in your list and retries the request there. Without the router, o4 only retries the same model. The router is off by default. Turn it on when you have keys for more than one provider and don’t want a busy provider to stop your work. The router keeps one model per provider, so you need credentials for at least two providers.

Set up the router

1

Preview the configuration

Run /router setup with the models you want to fall back to, in order:
Without --yes, o4 only shows the router status the configuration would give: the primary model, the usable fallbacks and any skipped models. Without --candidate, o4 picks the models itself (see Which models the router can use).If fewer than two providers are usable, you see Router setup was not enabled. with a hint on which key to add, and o4 writes nothing, even with --yes.
2

Write it

Add --yes to save it to ~/.o4/config.toml. Add --project to save it to the project’s .o4/config.toml instead:
o4 sets enabled = true and replaces the candidates list in that file with the models you pass. Without --candidate, it removes the list, so o4 goes back to picking the models itself. --capacity replaces the capacity limits; without it, the limits already in the file stay. Other [router] keys stay as they are. o4 doesn’t write if a file that takes precedence already has a [router] section: the project’s .o4/config.toml or .o4/config.local.toml, or ~/.o4/managed/config.toml. Edit that file instead.
3

Start a new session

Router changes take effect in the next session.
You can also write the [router] section by hand:

From the settings

The Providers tab of /config has three router rows:
  • Router setup shows the same preview as /router setup with no options. It never writes the configuration; run /router setup --yes for that. o4 shows this row while the router is off.
  • Router disable takes the place of Router setup while the router is on. It runs /router disable.
  • Router status runs /router status.

Which models the router can use

The first usable model in the list is the one the session starts on. o4 builds the list like this:
  • With candidates: the model you passed with -m, then the session’s saved model when you resume, then the models in candidates, in order.
  • Without candidates: the model you passed with -m, the session’s saved model, your default model, the default model of your default provider, then the default model of each other provider you have credentials for.
o4 leaves out any model it can’t find in the model catalog, and any model whose provider has no credentials. A built-in ollama model that points at a local server (localhost or 127.0.0.1) needs none. The list keeps only the first model of each provider; with candidates, /router status warns about each extra model from a provider that is already in the list. With the router on, a model you pass with -m that is missing or has no credentials doesn’t stop o4. The session starts on the next usable model instead.
Without candidates, “each other provider you have credentials for” means providers with a key stored in o4’s settings, custom providers whose key variable is set, and the kimi-coding, openai-codex and claude-code providers when o4 finds their credentials. A built-in provider whose key is only in an environment variable, such as ANTHROPIC_API_KEY, joins the list only as your default provider. If /router setup finds too few providers, store the keys under /config providers, or list the models in candidates, where environment variables count.
When a request fails, o4 also skips a model for this switch if:
  • its context window is unknown, or the conversation already fills more than context_safety_ratio of it (75% by default);
  • it has no tool support, or no reasoning support when the session started with a reasoning level other than default. In o4 0.2.74, the router doesn’t treat built-in ollama models or the Bedrock Kimi models as tool-capable, so it never switches to them;
  • it’s the model that just failed.

What happens when a request fails

  1. o4 sorts the error into one of four kinds: rate_limit (including HTTP 429), overloaded, timeout (including dropped connections), or server_error (HTTP 500, 502 and 503). Any other error, such as a bad API key, never causes a fallback.
  2. If that kind is in fallback_on, o4 switches to the next usable model in list order and retries the request on it. A notice and the activity row show Router switched route (<from> -> <to>) with the reason.
  3. If no model is usable, or the kind isn’t in fallback_on, the activity row shows Router fallback unavailable with the reason, and o4 falls back to its normal retries on the current model.
  4. After max_fallback_attempts switches (1 by default), o4 stops switching and ends the turn with the error. The count starts again after a turn that succeeds.
The router acts only on errors that o4 would retry anyway, and only before the model has streamed any text in that reply. If a reply fails partway through, the turn ends with the error. After a switch, the session stays on the new model. The router doesn’t switch back on its own. To go back, pick the model with /model. A model you pick by hand becomes the active model, and the router keeps the rest of the list as fallbacks.

Limit requests per provider

You can cap how many requests the session sends to one provider at the same time. The cap is shared by all tabs of the interactive interface, and applies to print mode runs too. It doesn’t cover subagents or campaign workers, and it only works while the router is on:
Or with /router setup:
Extra requests wait until one finishes.

Check the router

/router on its own does the same. It shows whether the router is on, whether a fallback is active, the primary model, the number of providers it can route to, each usable fallback with its context window, each skipped model with the reason (model not found or missing auth) and a hint on which key to add, any capacity limits, and warnings such as a second model from the same provider. Fallback is active only when the router is on and at least two providers are usable. /router status reads the [router] settings the session started with, and your current keys. It shows the list o4 would build now, which can differ from the session’s running list, for example after a fallback or a /model switch.

Turn the router off

This sets enabled = false in ~/.o4/config.toml and keeps the rest of the [router] section. Add --project to change the project’s .o4/config.toml. /router off does the same. The change takes effect in the next session. If a file that takes precedence sets enabled = true, the router stays on; set it there instead.

Settings

All [router] settings, with their types and defaults, are in the configuration reference: o4 checks these values when a session starts. It replaces an out-of-range max_fallback_attempts or context_safety_ratio with the default, drops unknown fallback_on entries and capacity limits below 1, and logs a warning for each.