Skip to main content
This guide shows how to use the Gloo AI Responses API. You’ll build four working examples: a basic guarded response with a theological tradition, instructions with multi-turn input, streaming, and vision (image input).
Why Responses? The Responses API is the recommended way to build on Gloo AI. It uses the OpenAI-compatible Responses request shape and is guarded by default — every request runs the same guardrails, values-aligned tradition responses, and output moderation that power Completions V2, with no extra configuration.

Prerequisites

Before starting, ensure you have:

Understanding the Responses API

All requests go to a single guarded endpoint: POST /ai/v2/guarded/responses

Key Request Fields

Answers come back as a typed output[] array. See Example 1’s “Understanding the Response” for the full shape and how to extract the text.

Guardrail Behavior

Values-aligned static answers arrive as normal messages. When guardrails intervene with a curated, tradition-appropriate answer, it comes back as a regular message item in output[] — identical shape to a model answer. Your client needs no special handling; just extract the text as usual. Hard blocks return 403. Requests that guardrails reject outright return an HTTP 403:
A 403 is not a transient failure — don’t blindly retry. Adjust the request content instead. Model text output also passes through Gloo’s output-moderation layer before it reaches you — for the full pipeline, see Guardrails and values in the Responses API Guide. See the Responses API Guide for complete endpoint documentation.

Moving from Completions

If you already use Completions V2, the migration is mostly a rename:
Auto-routing and model_family selection remain Completions V2 features. The Responses API is pinned-model only — you specify an exact model on every request.

Example 1: Basic Guarded Response with Tradition

Send a plain-string input and a tradition to get a values-aligned answer:

Understanding the Response

The answer comes back as a typed output[] array — a message item whose content[] holds output_text parts:
To extract the text, walk output[] for message items and join their output_text parts — the cookbook samples do exactly this in an extractText helper.

What You’ll See

Running the full sample prints one block per example, then the final banner:

Example 2: Instructions + Multi-Turn Input

Use instructions for system-level guidance and pass the conversation as an input array. The history alternates user → assistant → user (three turns), so the model continues the conversation with full context:

What You’ll See

The reply arrives as the same message item in output[] shown in Example 1, continuing the conversation. Running the full sample prints the second block:

Example 3: Streaming (SSE)

Set "stream": true to receive the response as Server-Sent Events.

Understanding the Response

When you set "stream": true, the API response is an SSE stream. Here are the key event types, and what to do with each: On the wire, each event is an event: line followed by a data: line:
The live API sends flattened payloads: each data: line is the event object itself (its type matches the event: line), with no nested response wrapper. The response.completed payload carries usage at the top level and does not include the output[] array — the full text arrives via response.output_text.done. The snippets below accept both the flattened shape and the nested shape so they work with either.

What You’ll See

When you run the full sample, the deltas print live as they arrive. After the stream completes, the runner prints the model, the streamed-text length with a preview, and the usage from response.completed:
Streams can fail after the response has started — a dropped connection mid-stream means partial output. See Handling Streaming Failures for retry, continuation, and partial-output guidance.

Example 4: Vision (Image Input)

Vision uses the same endpoint — pass typed content parts (input_text + input_image) in the content of an input item. Use a vision-capable model such as gloo-google-gemini-3.1-pro:

What You’ll See

The reply arrives as a normal message item in output[] — for the cat photo above, the answer identifies the animal:
Running the full sample prints the fourth block:

Complete Example

View Complete Code

Clone or browse the complete working examples for all 6 languages (JavaScript, TypeScript, Python, PHP, Go, Java) with setup instructions.

What’s Included

Each language implementation provides:
  • auth — Shared API key configuration
  • config — Centralized configuration (endpoint, env vars)
  • Runner per example — Basic guarded response, instructions + multi-turn, streaming, and vision, plus a final banner
  • Env helper — Ignores unset .env placeholders
  • post helper — Shared request helper that annotates 403s with “(guardrails hard-blocked this request)”
  • extractText helper — Reusable helper for pulling text out of output[]
Set up your credentials first:
Then run per language:
Each runner prints a block per example (model used, first ~100 characters of the response, token usage, and a ✓ pass line), then the final banner:
The snippets in this tutorial are simplified to be self-contained. The cookbook samples are structured the same way but add robustness: an env helper that ignores unset .env placeholders, a shared post helper that annotates 403s with “(guardrails hard-blocked this request)”, and a reusable extractText helper for pulling text out of output[]. The prompts, field names, and endpoint are identical.

Troubleshooting

Please set your GLOO_API_KEY environment variable : Your .env file is missing or doesn’t contain GLOO_API_KEY. Create one in the language’s directory (see Complete Example), or export the variable in your shell for Go and Java. API request failed with status 403 : Guardrails hard-blocked the request. The 403 body includes a content_policy_violation detail — rephrase the request rather than retrying. See Guardrail behavior. Vision request returns 400 / INVALID_REQUEST : The model isn’t vision-capable. Vision requires a multimodal model like gloo-google-gemini-3.1-pro — sending image parts to a text-only model (e.g. gloo-anthropic-claude-sonnet-4.6) fails. Check Supported Models for capabilities. Vision request fails on image fetch : The API fetches image_url server-side, so the URL must be publicly reachable. Test it in a browser or curl -I <url>. Private or expiring URLs fail; use a public URL or a base64 data: URI. Stream parses nothing, or every event looks unknown : Remember the payloads are flattened — the data: line is the event object itself, not { "response": {...} }. Match on the event: line, handle response.output_text.delta/done and response.completed, and ignore other response.* lifecycle events. Treating unknown events as terminal (and closing the stream) is the safe default.

Next Steps

Now that you understand the Responses API, explore:
  1. Responses API Guide - Full API documentation
  2. Tool Use - Function calling with Responses
  3. Grounded Responses - RAG with source attribution
  4. Direct Responses - Unguarded direct endpoint
  5. Completions V2 Guide - Auto-routing and model_family selection
  6. Supported Model IDs - All available models