用 Responses 生成
POST /v1/responses
OpenAI Responses 兼容的生成接口。传 model 和 input,input 可以是字符串或受支持的结构化输入项。生成的文本要从 output[] 里的消息内容读,不是某个统一的顶层文本字段;逐项判断类型,不要假定 output[0] 就是文本。加 stream: true 返回 SSE,通常是 response.output_text.delta 和 response.completed。每个接口都要求模型的 supported_endpoints 精确包含该路径。这几个接口转发所选模型自己的协议,可选参数和响应细节会随本地 worker 或远程供应方不同。本地模型把输入留在设备上;配置为远程的模型会收到发给它的请求。
请求体
application/json必填
modelstring必填文本模型 ID,从模型清单里取。它的 supported_endpoints 必须包含这个路径。input必填字符串,或受支持的结构化输入项。max_output_tokensinteger生成 token 的上限。不传则用模型自己的上限。streambooleantrue 返回 SSE 事件流,而不是单个 JSON 响应。响应
200生成的响应;stream: true 时是 SSE 流。生成文本、token 数和 ID 每次都不同。idstringobjectstringstatusstringmodelstringoutputobject[]typestringrolestringstatusstringcontentobject[]{
"id": "resp_example",
"object": "response",
"status": "completed",
"model": "Qwen/Qwen3.5-4B",
"output": [
{
"type": "message",
"role": "assistant",
"status": "completed",
"content": [
{
"type": "output_text",
"text": "Hello!"
}
]
}
]
}400请求不合法。先改请求再重试:error.code 用于程序判断,error.param 用于定位是哪个输入。401API Key 缺失或无效。403Host 或 Origin 不被允许,或 License 被拒。按 error.code 区分这两种情况。413请求体超过 512 MiB。503所需的本地模型仍在准备(model_downloading),或服务正忙(service_busy)。带 Retry-After 时按它退避。