> ## Documentation Index
> Fetch the complete documentation index at: https://docs.winterr.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# SDK inference contract

`GET /v1/models` and `POST /v1/responses` are marked `contract-only` in the API
contract. The SDKs implement these request and response shapes. This guide does
not establish a hosted inference service. Use the methods only in an environment
that explicitly supports them.

## Requests and results

Both routes require bearer authentication with `inference:invoke` scope for the
project. `models.list()` returns a model list with a `data` array. Check that the
array contains the model you need before you select its ID.

A response request requires a nonempty `model` and `input`. Input is a nonempty
text string, or a nonempty array of messages. Each message has a `developer`,
`system`, or `user` role and a nonempty content array of
`{"type": "input_text", "text": "Hello"}` objects.

Optional fields are `instructions`, `max_output_tokens`, `metadata`, `stream`,
`temperature`, and `top_p`. A non-null output token limit must be a positive
integer. Temperature is from 0 to 2; top-p is from 0 to 1. Metadata values are
strings with at most 512 characters each. Additional request fields are outside
this contract. Tools, image input, audio input, embeddings, and Chat Completions
are not part of this typed Responses contract.

A non-streaming call returns the response object directly, not a resource and
operation pair. Inspect its status, output, error, incomplete details, and usage.
Do not treat HTTP success alone as a completed model response.

## Client methods

| Task                                 | TypeScript                                                             | Python                               |
| ------------------------------------ | ---------------------------------------------------------------------- | ------------------------------------ |
| List models                          | `client.models.list()`                                                 | `client.models.list()`               |
| Create a response                    | `client.responses.create(input)`                                       | `client.responses.create(input)`     |
| Read events                          | `client.responses.stream(input)` or `create({...input, stream: true})` | `client.responses.stream(input)`     |
| Read response headers and request ID | `client.responses.createRaw(input)`                                    | `client.responses.create_raw(input)` |

Python `create()` rejects `stream: true`; use `stream()` instead. `AsyncWinterr`
provides asynchronous methods and asynchronous iteration. TypeScript streams use
an async iterable. Events contain the response event directly. Read text from
`response.output_text.delta` events and check the terminal event.

Both clients stop after the first `response.completed`, `response.incomplete`,
`response.failed`, or `error` event. An incomplete or failed event is not success.
A connection that ends without a terminal event does not prove completion. Close
the iterator when you stop reading; use `aclose()` for a Python async iterator.
Closing a stream releases client resources but does not prove server cancellation.
The contract defines no response cancel or drain route.

## Retries and errors

Response creation is not idempotent. Both SDKs disable automatic retries and do
not generate an idempotency key for it. Do not apply the resource-create retry
rule or add an idempotency header to make this request retryable. Repeating a
request after a lost response can submit new work.

Use `validation: "strict"` in TypeScript or `validation="strict"` in Python to
check known response shapes locally. This does not prove model correctness or
service availability. API errors retain the API code, HTTP status, request ID,
and details. Keep the request ID for support; do not include credentials or
private prompts in a support log.

## Service availability

Models and Responses remain contract-only. Hosted model execution, live response
streaming and Chat Completions are not available through these SDK methods.
Request validation does not grant access to a model or reserve inference funds.

Live service use still requires model selection, trusted service and runtime
identity, production token metering, remote cancellation and release checks.
Client stream cleanup does not prove that remote execution stopped. Use this
contract only in an environment that explicitly supports the required methods.
