GET /v1/models and POST /v1/responses are marked contract-only in the API
contract. The SDKs implement these request and response shapes. This guide does
not establish a hosted inference service. Use the methods only in an environment
that explicitly supports them.
Requests and results
Both routes require bearer authentication withinference:invoke scope for the
project. models.list() returns a model list with a data array. Check that the
array contains the model you need before you select its ID.
A response request requires a nonempty model and input. Input is a nonempty
text string, or a nonempty array of messages. Each message has a developer,
system, or user role and a nonempty content array of
{"type": "input_text", "text": "Hello"} objects.
Optional fields are instructions, max_output_tokens, metadata, stream,
temperature, and top_p. A non-null output token limit must be a positive
integer. Temperature is from 0 to 2; top-p is from 0 to 1. Metadata values are
strings with at most 512 characters each. Additional request fields are outside
this contract. Tools, image input, audio input, embeddings, and Chat Completions
are not part of this typed Responses contract.
A non-streaming call returns the response object directly, not a resource and
operation pair. Inspect its status, output, error, incomplete details, and usage.
Do not treat HTTP success alone as a completed model response.
Client methods
Python
create() rejects stream: true; use stream() instead. AsyncWinterr
provides asynchronous methods and asynchronous iteration. TypeScript streams use
an async iterable. Events contain the response event directly. Read text from
response.output_text.delta events and check the terminal event.
Both clients stop after the first response.completed, response.incomplete,
response.failed, or error event. An incomplete or failed event is not success.
A connection that ends without a terminal event does not prove completion. Close
the iterator when you stop reading; use aclose() for a Python async iterator.
Closing a stream releases client resources but does not prove server cancellation.
The contract defines no response cancel or drain route.
Retries and errors
Response creation is not idempotent. Both SDKs disable automatic retries and do not generate an idempotency key for it. Do not apply the resource-create retry rule or add an idempotency header to make this request retryable. Repeating a request after a lost response can submit new work. Usevalidation: "strict" in TypeScript or validation="strict" in Python to
check known response shapes locally. This does not prove model correctness or
service availability. API errors retain the API code, HTTP status, request ID,
and details. Keep the request ID for support; do not include credentials or
private prompts in a support log.