Skip to main content
Direct Responses is the direct endpoint for the Responses format: POST /ai/v2/direct/responses sends your request straight to the model you name. It exposes the entire Gloo AI model catalog through a single endpoint — text, vision, reasoning, tool use, and native image generation — with one request and response shape that works the same across every provider. It takes the same request body as the guarded Responses endpoint; the difference is what runs around it. Nothing here applies guardrails, output moderation, values-alignment (tradition), model_family selection or intelligent routing. Choose this when your application owns its own safety and prompting; choose guarded Responses — the recommended default — when it should inherit Gloo’s. See Which endpoint should I use? for the full comparison across all three endpoint families.
New to Gloo AI? Start on the guarded Responses endpoint — same request shape, with Gloo’s guardrails and values-alignment applied — and come here only if you need the unprocessed path. If you already use Completions V2, it remains fully supported and backwards-compatible — see Moving from Completions to Responses.

Why the Responses API

The Responses API offers a more capable, forward-looking request shape that the broader ecosystem is standardizing on:
  • One shape for every modality. Text, image input (vision), and image generation all use the same input array and the same typed output array — no separate endpoints or bespoke payloads per capability.
  • OpenAI-compatible. The request and response formats mirror the OpenAI Responses API, so existing tooling, SDKs, and mental models carry over directly. Point your base URL at https://platform.ai.gloo.com/ai/v2/direct and go.
  • Typed, structured output. Responses come back as an output[] array of typed items (message, image_generation_call, reasoning, tool calls) instead of a single opaque choices[].message.content string — easier to parse, and extensible as new item types arrive.
  • Built for multimodal. Native image generation and image input are first-class, not bolted on. The same endpoint that answers a text prompt can return a generated image.
If you are starting a new integration, build on the Responses API.

Endpoint

URL: https://platform.ai.gloo.com/ai/v2/direct/responses Operation: POST
Authentication is identical to the rest of the platform — send your API key as a Bearer token. See Generate API Keys.

Request format

The Responses API uses input instead of messages, and a few renamed fields. If you know the chat-completions format, the mapping is small:
Conversation state is managed client-side by passing the full input history on each request. The previous_response_id parameter (used by OpenAI’s Responses API for server-side conversation chaining) is not supported.
This is a direct endpoint: /ai/v2/direct/responses does not apply the governance layer. Requests here call the exact model you specify with no additional processing. None of the following runs:
  • Input guardrails and output moderation
  • Values-alignment (tradition)
  • model_family selection and intelligent auto-routing
Want the same request body with those applied? Use the guarded Responses endpoint at /ai/v2/guarded/responses — same body, different URL. Completions V2 and Grounded Completions are guarded too, on the Chat Completions shape.“Direct” describes this URL, not the Responses format in general: the same format is served by the guarded endpoint above and by Grounded Responses, which adds retrieval-augmented (RAG) generation over your own content.

Response format

Responses return a typed output[] array. Each item has a type; a normal text answer arrives as a message item, a generated image as an image_generation_call item.

Streaming

Set "stream": true to receive the response as Server-Sent Events. Each event has an event: line and a data: line carrying a typed JSON payload:
Key event types to handle:
  • response.output_item.added / response.output_item.done — mark the start and end of each typed item in output[] (including message, image_generation_call, and function_call).
  • response.content_part.added / response.content_part.done — mark the start and completion of a content part within an item. response.content_part.done is the signal that a message’s text is complete, followed by response.output_item.done.
  • response.output_text.delta — incremental text tokens for a message item.
  • response.function_call_arguments.delta — incremental tool-call argument tokens for a function_call item.
  • response.completed — emitted once when the response is finished; the final payload includes the full output[] array and usage.
Errors are not delivered as SSE events. They surface as an HTTP error status (or a raised exception in the SDK), so handle them on the request itself rather than watching for an error event in the stream.

Multimodal

Image input (vision)

Pass images as input items alongside text. Any vision-capable model accepts them:
Images may be supplied as a remote URL or a base64 data: URI.

Image generation

Image-capable models return a generated image as an image_generation_call output item (base64 result). Optional image_generation controls (quality, size) are forwarded to providers that support them.
The generated image comes back as a typed image_generation_call item in output[]. The result field carries the image as a base64-encoded string:
Decode the base64 result to get the image bytes. Optional image_generation controls (quality, size) are forwarded to providers that support them.

OpenAI: image generation via the image_generation tool

OpenAI’s image generation runs as an image_generation tool on a GPT-5 model — gpt-image-1 is not a top-level Responses model (exactly as on OpenAI’s own API, where it is an Images API model rather than a Responses one). Call a GPT-5 model and attach the tool; the model decides when to draw and can combine the image with text in the same response. The tool accepts an optional model field to pin its backing image model (e.g. gpt-image-1) plus quality/size controls:
The generated image arrives as the same image_generation_call output item shown above (alongside any reasoning and message items from the GPT-5 model). Image editing works through the same mechanism: include an input_image part in the conversation and it becomes the tool’s edit source. See Supported Models for which models support image input and image generation.

Pricing & token spend

The Responses API is billed per token at each model’s standard rates — there is no separate or premium price for using /responses. Cost is driven entirely by the model you select and the tokens you consume:
  • Per-model rates. Input and output rates vary by model. The live rates are on the Supported Models page and programmatically on GET /platform/v2/models.
  • Prompt caching reduces input cost when prompt prefixes repeat — cached tokens are billed at a discounted cache_read rate. See Prompt Caching.
  • Image generation is billed using the image model’s token accounting; check the model’s rates on the models endpoint.
  • A 6.5% Studio markup applies to every token segment, consistent with the rest of the platform.
Track real spend in the Gloo Studio billing dashboard and API usage.

Supported models

The Responses API works across the full Gloo AI catalog — Anthropic, OpenAI, Google, and open-source families — including the multimodal and image-generation models. Use the Model ID as the model field. The complete, live list (with capabilities and pricing) is on the Supported Models page.

Moving from Completions to Responses

Completions V2 (/ai/v2/guarded/chat/completions) is fully supported and backwards-compatible — existing integrations continue to work unchanged, and it remains the home of auto-routing and model_family selection. Guardrails, output moderation and tradition have already reached the Responses shape on the guarded Responses endpoint, and RAG grounding on Grounded Responses. The Responses API is the recommended surface for new work because it standardizes on the OpenAI-compatible Responses shape and makes multimodal (vision + image generation) first-class. Below the URL, the migration is mostly a rename:
The URL is the one row that is not a rename. Completions V2 applies guardrails, output moderation and tradition; this page’s endpoint applies none of them. POST /ai/v2/guarded/responses is the like-for-like destination — it keeps the governance layer while changing the request shape. Move to POST /ai/v2/direct/responses only if dropping that layer is a decision you are making deliberately, and the body stays identical either way.
tradition, guardrails and output moderation are available now on the guarded Responses endpoint, and grounded (RAG) responses on Grounded Responses. Auto-routing and model_family selection remain on Completions V2, and grounding with citation metadata on Grounded Completions. Choose this page’s direct endpoint when your application owns its own safety and prompting.