model in the response is the model you asked for. It is separate from auto-routing and model_family, which choose a model for you on the guarded Completions endpoint.
When a request fails over
Failover handles failures that belong to the provider, not to your request:- Provider errors and outages. A timeout, a lost connection, a provider server error, an empty response, or the provider refusing the request because it is over capacity or rate-limited. Gloo retries the request on another provider serving the same model.
- A provider that keeps failing. Gloo tracks each provider’s health per model. When a provider starts failing for a model, new requests for that model go to a healthy provider instead of trying the failing one first. It is brought back into service automatically once it recovers.
- A stream that fails before the answer starts. If a streamed request fails before any text reaches you, Gloo retries it, on another provider when needed, and the answer starts from the beginning. You see the stream start later, not an error.
- A stream cut off mid-answer. For models that support it, another provider continues the answer from where the stream stopped, in the same response. For models that do not, the stream ends with an error event instead; see Handling Streaming Failures for how to continue it from your application.
Not every model has a second provider. For a model served from only one place, failover has nowhere to send the request, and a provider outage returns an error.
When a request does not fail over
Some failures would fail the same way on every provider, so Gloo returns them straight away instead of repeating the request elsewhere:Reasoning models and the output limit
Models that think before they answer produce their reasoning inside the same output-token budget as their visible answer. If you set a smallmax_tokens (Chat Completions) or max_output_tokens (Responses), the model can spend all of it reasoning and stop before writing any text.
The provider answered correctly and would answer the same way again, so this is not retried or failed over. Gloo returns the provider’s own answer:
- Chat Completions, streaming and non-streaming: the response ends with
finish_reason: "length"and no text, andusagereports the tokens the model used. - Responses (non-streaming, OpenAI reasoning models): HTTP
200withstatus: "incomplete"andincomplete_details: {"reason": "max_output_tokens"}, plus asuggestionfield that names the tokens used and the limit you sent.
Which endpoints fail over
What you observe
- The same model. The response carries the
modelyou requested, whichever provider served it. - Extra latency. A request that failed over waited for the failed attempt first, so it takes longer than usual.
- A colder prompt cache. Prompt caches are held by each provider, so a request served by a different provider does not read the cache built on the original one. Expect fewer cached tokens on requests that failed over, and on requests routed around an unhealthy provider, until the cache warms on the new one. See Prompt Caching.
- No charge for the failed attempt. An attempt that failed and was replaced by another provider is not billed.
- An error only when failover could not help. Errors you receive carry the standard error object:
faultsays whether the problem was your request, the provider, or Gloo, andretryablesays whether the same request may succeed if you send it again.
What your application should do
- Rely on failover for provider trouble. You do not need to switch models or add a second provider yourself to ride out a provider outage.
- Still retry retryable errors. When an error has
retryable: true, retry with exponential backoff. A later request benefits from Gloo having already routed around the failing provider. Do not retry errors withretryable: falseunchanged. - Handle interrupted streams. On models where Gloo cannot continue a cut-off stream, the stream ends with an error event after partial output. Follow Handling Streaming Failures.
- Size output limits for reasoning models. Treat
finish_reason: "length"orstatus: "incomplete"with no text as a sign to raisemax_tokensormax_output_tokens, not as an outage. - Log
trace_id. Include it when you contact support about a failed request.
Related Documentation
- Endpoint Types — what each endpoint family runs around your request
- Errors — error codes,
fault, andretryable - Handling Streaming Failures — retry and continuation after a stream fails
- Prompt Caching — how caching works on each provider

