用 Responses 生成

POST /v1/responses

OpenAI Responses 兼容的生成接口。传 model 和 input,input 可以是字符串或受支持的结构化输入项。生成的文本要从 output[] 里的消息内容读,不是某个统一的顶层文本字段;逐项判断类型,不要假定 output[0] 就是文本。加 stream: true 返回 SSE,通常是 response.output_text.delta 和 response.completed。每个接口都要求模型的 supported_endpoints 精确包含该路径。这几个接口转发所选模型自己的协议,可选参数和响应细节会随本地 worker 或远程供应方不同。本地模型把输入留在设备上;配置为远程的模型会收到发给它的请求。

指南与示例语言与视觉全部接口

请求体

application/json必填

modelstring必填文本模型 ID,从模型清单里取。它的 supported_endpoints 必须包含这个路径。
input必填字符串,或受支持的结构化输入项。
max_output_tokensinteger生成 token 的上限。不传则用模型自己的上限。
streambooleantrue 返回 SSE 事件流,而不是单个 JSON 响应。

响应

200生成的响应;stream: true 时是 SSE 流。生成文本、token 数和 ID 每次都不同。
idstring
objectstringresponse
statusstring
modelstring
outputobject[]
typestring
rolestring
statusstring
contentobject[]
JSON
{
  "id": "resp_example",
  "object": "response",
  "status": "completed",
  "model": "Qwen/Qwen3.5-4B",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "status": "completed",
      "content": [
        {
          "type": "output_text",
          "text": "Hello!"
        }
      ]
    }
  ]
}
400请求不合法。先改请求再重试:error.code 用于程序判断,error.param 用于定位是哪个输入。
401API Key 缺失或无效。
403Host 或 Origin 不被允许,或 License 被拒。按 error.code 区分这两种情况。
413请求体超过 512 MiB。
503所需的本地模型仍在准备(model_downloading),或服务正忙(service_busy)。带 Retry-After 时按它退避。