Skip to main content
GET /v1/models and POST /v1/responses are marked contract-only in the API contract. The SDKs implement these request and response shapes. This guide does not establish a hosted inference service. Use the methods only in an environment that explicitly supports them.

Requests and results

Both routes require bearer authentication with inference:invoke scope for the project. models.list() returns a model list with a data array. Check that the array contains the model you need before you select its ID. A response request requires a nonempty model and input. Input is a nonempty text string, or a nonempty array of messages. Each message has a developer, system, or user role and a nonempty content array of {"type": "input_text", "text": "Hello"} objects. Optional fields are instructions, max_output_tokens, metadata, stream, temperature, and top_p. A non-null output token limit must be a positive integer. Temperature is from 0 to 2; top-p is from 0 to 1. Metadata values are strings with at most 512 characters each. Additional request fields are outside this contract. Tools, image input, audio input, embeddings, and Chat Completions are not part of this typed Responses contract. A non-streaming call returns the response object directly, not a resource and operation pair. Inspect its status, output, error, incomplete details, and usage. Do not treat HTTP success alone as a completed model response.

Client methods

Python create() rejects stream: true; use stream() instead. AsyncWinterr provides asynchronous methods and asynchronous iteration. TypeScript streams use an async iterable. Events contain the response event directly. Read text from response.output_text.delta events and check the terminal event. Both clients stop after the first response.completed, response.incomplete, response.failed, or error event. An incomplete or failed event is not success. A connection that ends without a terminal event does not prove completion. Close the iterator when you stop reading; use aclose() for a Python async iterator. Closing a stream releases client resources but does not prove server cancellation. The contract defines no response cancel or drain route.

Retries and errors

Response creation is not idempotent. Both SDKs disable automatic retries and do not generate an idempotency key for it. Do not apply the resource-create retry rule or add an idempotency header to make this request retryable. Repeating a request after a lost response can submit new work. Use validation: "strict" in TypeScript or validation="strict" in Python to check known response shapes locally. This does not prove model correctness or service availability. API errors retain the API code, HTTP status, request ID, and details. Keep the request ID for support; do not include credentials or private prompts in a support log.

Service availability

Models and Responses remain contract-only. Hosted model execution, live response streaming and Chat Completions are not available through these SDK methods. Request validation does not grant access to a model or reserve inference funds. Live service use still requires model selection, trusted service and runtime identity, production token metering, remote cancellation and release checks. Client stream cleanup does not prove that remote execution stopped. Use this contract only in an environment that explicitly supports the required methods.