Skip to main content
The Gloo AI platform provides access to a wide range of leading models from multiple providers. Visit the Model Explorer in Gloo Studio for a richer side-by-side comparison including reasoning capabilities, modalities, and speed ratings. The table below is fetched live in your browser from the public, unauthenticated GET /platform/v2/models endpoint — so it always reflects the current platform catalog. Use the Model ID as the model parameter on the endpoint for that model type. Generation models work with Responses and Completions; models marked Embedding work with the Embeddings API.
The rates shown are list prices. Guarded generation endpoints add the platform rate described in Studio Billing; direct endpoints, including Embeddings, bill at the published rate. Track real spend in the Gloo Studio billing dashboard.
Generation models. Gloo AI can route generation requests to providers that serve the model. Embedding requests use the embedding model you name and are not failed over to another provider. See Provider Failover.
Generation model IDs work across the Responses API and Completions V2. Embedding model IDs are for the Embeddings API. Capability flags (supports_tools, supports_streaming, supports_reasoning, supports_vision) apply to generation; the model catalog returns them for all entries.

Deprecated Models

Models being retired are flagged live in the catalog table below: a deprecated model renders a Deprecated badge next to its ID, with its deprecation note inline. Those come from the is_deprecated and deprecation_note fields on GET /platform/v2/models, alongside replacement_model — the ID that requests are routed to. For generation models, where a replacement has been chosen, requests using the deprecated ID may continue to be routed to that replacement while you migrate. Do not assume replacement routing applies to embeddings: changing embedding models requires re-embedding stored content. Read deprecation_note for the specifics, including any price change or missing replacement. See Model Lifecycle for how a replacement is chosen, what the price can do, how much notice you get, and how to detect all of it from the API.

Prompt Caching

For generation models, the Caching column shows which models support prompt caching, and which type — Implicit (automatic, e.g. OpenAI, DeepSeek, Gemini, Qwen) or Explicit (opt-in per request, e.g. Anthropic). Caching does not apply to embedding models. For full details on each provider’s caching mechanism, billing rates, and best practices, see the dedicated Prompt Caching guide.