Generate with Responses

POST /v1/responses

OpenAI Responses-compatible generation. Send model and input, which can be a string or supported structured input items. Read generated text from message content inside output[] rather than a universal top-level text field, and inspect item types instead of assuming output[0] is text. Adding stream: true returns SSE, typically response.output_text.delta and response.completed. Each endpoint requires a model whose supported_endpoints includes that exact path. These APIs forward the selected model's protocol, so optional parameters and response details can vary by local worker or remote provider. Local models keep inputs on the device; a configured remote model receives the request sent to it.

Guide and examplesLanguage and visionAll endpoints

Request body

application/jsonRequired

modelstringRequiredText model ID from the model catalog. Its supported_endpoints must contain this path.
inputRequiredA string, or supported structured input items.
max_output_tokensintegerUpper bound on generated tokens. Omit it to use the model's own limit.
streambooleantrue returns Server-Sent Events instead of a single JSON body.

Responses

200The generated response, or an SSE stream when stream: true. Generated text, token counts and IDs vary.
idstring
objectstringresponse
statusstring
modelstring
outputobject[]
typestring
rolestring
statusstring
contentobject[]
JSON
{
  "id": "resp_example",
  "object": "response",
  "status": "completed",
  "model": "Qwen/Qwen3.5-4B",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "status": "completed",
      "content": [
        {
          "type": "output_text",
          "text": "Hello!"
        }
      ]
    }
  ]
}
400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.
401Invalid or missing API key.
403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.
413Request body exceeds 512 MiB.
503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.