> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gloo.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Provider Failover

> How Gloo AI keeps serving the model you asked for when the provider behind it fails, when it deliberately does not retry, and what your application should still handle.

Most models in the Gloo AI catalog can be served from more than one provider. When the provider currently serving a model is failing, Gloo sends your request to another provider that serves **the same model**. This is **provider failover**, and it runs on every chat and responses endpoint without any configuration.

Failover changes *who* serves your request, never *what* you get. It never substitutes a different model: the `model` in the response is the `model` you asked for. It is separate from [auto-routing and `model_family`](/api-guides/completions-v2), which choose a model for you on the guarded Completions endpoint.

## When a request fails over

Failover handles failures that belong to the provider, not to your request:

* **Provider errors and outages.** A timeout, a lost connection, a provider server error, an empty response, or the provider refusing the request because it is over capacity or rate-limited. Gloo retries the request on another provider serving the same model.
* **A provider that keeps failing.** Gloo tracks each provider's health per model. When a provider starts failing for a model, new requests for that model go to a healthy provider instead of trying the failing one first. It is brought back into service automatically once it recovers.
* **A stream that fails before the answer starts.** If a streamed request fails before any text reaches you, Gloo retries it, on another provider when needed, and the answer starts from the beginning. You see the stream start later, not an error.
* **A stream cut off mid-answer.** For models that support it, another provider continues the answer from where the stream stopped, in the same response. For models that do not, the stream ends with an error event instead; see [Handling Streaming Failures](/best-practices/completions-streaming-failures) for how to continue it from your application.

Failover works within a time budget for each request. If no provider answers within it, you receive the error rather than a request that waits indefinitely.

<Note>
  Not every model has a second provider. For a model served from only one place, failover has nowhere to send the request, and a provider outage returns an error.
</Note>

## When a request does not fail over

Some failures would fail the same way on every provider, so Gloo returns them straight away instead of repeating the request elsewhere:

| Situation | What you get | What to do |
| :- | :- | :- |
| Invalid request, unsupported model, or input longer than the model's context window | A `400` client error | Fix the request. See [Errors](/api-reference/general/errors). |
| Your API key or organization is not authorized | A `401` or `403` | Check your credentials. |
| The provider's content policy blocks the request or the answer | A content-filter error, or `finish_reason: "content_filter"` on a stream | Change the request. Do not retry it unchanged. |
| Your organization is out of credit or at its spending limit | A `402` or `429` with `fault: "client"` | Resolve the limit. See [Limits](/api-reference/general/limits). |
| A reasoning model spent your whole output limit before writing any visible text | A normal response that stopped at the limit (below) | Raise the output limit. |

### Reasoning models and the output limit

Models that think before they answer produce their reasoning inside the same output-token budget as their visible answer. If you set a small `max_tokens` (Chat Completions) or `max_output_tokens` (Responses), the model can spend all of it reasoning and stop before writing any text.

The provider answered correctly and would answer the same way again, so this is **not** retried or failed over. Gloo returns the provider's own answer:

* **Chat Completions**, streaming and non-streaming: the response ends with `finish_reason: "length"` and no text, and `usage` reports the tokens the model used.
* **Responses** (non-streaming, OpenAI reasoning models): HTTP `200` with `status: "incomplete"` and `incomplete_details: {"reason": "max_output_tokens"}`, plus a `suggestion` field that names the tokens used and the limit you sent.

The fix is to raise the limit. Leave generous headroom above the length of the answer you expect.

## Which endpoints fail over

| Endpoint | Failover |
| :- | :- |
| Guarded [Responses](/api-guides/responses) and [Completions](/api-guides/completions-v2) | Yes |
| [Direct Responses](/api-guides/direct-responses) and [Direct Completions](/api-guides/direct-completions) | Yes |
| [Grounded Responses](/api-guides/grounded-responses) and [Grounded Completions](/api-guides/grounded-completions) | Yes |
| [Embeddings](/api-guides/embeddings) | No. A provider failure returns a `503` with `retryable: true`. |

## What you observe

* **The same model.** The response carries the `model` you requested, whichever provider served it.
* **Extra latency.** A request that failed over waited for the failed attempt first, so it takes longer than usual.
* **A colder prompt cache.** Prompt caches are held by each provider, so a request served by a different provider does not read the cache built on the original one. Expect fewer cached tokens on requests that failed over, and on requests routed around an unhealthy provider, until the cache warms on the new one. See [Prompt Caching](/api-guides/prompt-caching).
* **No charge for the failed attempt.** An attempt that failed and was replaced by another provider is not billed.
* **An error only when failover could not help.** Errors you receive carry the standard [error object](/api-reference/general/errors#ai-error-object): `fault` says whether the problem was your request, the provider, or Gloo, and `retryable` says whether the same request may succeed if you send it again.

## What your application should do

1. **Rely on failover for provider trouble.** You do not need to switch models or add a second provider yourself to ride out a provider outage.
2. **Still retry retryable errors.** When an error has `retryable: true`, retry with exponential backoff. A later request benefits from Gloo having already routed around the failing provider. Do not retry errors with `retryable: false` unchanged.
3. **Handle interrupted streams.** On models where Gloo cannot continue a cut-off stream, the stream ends with an error event after partial output. Follow [Handling Streaming Failures](/best-practices/completions-streaming-failures).
4. **Size output limits for reasoning models.** Treat `finish_reason: "length"` or `status: "incomplete"` with no text as a sign to raise `max_tokens` or `max_output_tokens`, not as an outage.
5. **Log `trace_id`.** Include it when you contact support about a failed request.

## Related Documentation

* [Endpoint Types](/api-guides/endpoint-types) — what each endpoint family runs around your request
* [Errors](/api-reference/general/errors) — error codes, `fault`, and `retryable`
* [Handling Streaming Failures](/best-practices/completions-streaming-failures) — retry and continuation after a stream fails
* [Prompt Caching](/api-guides/prompt-caching) — how caching works on each provider
