Generate with Chat Completions

POST /v1/chat/completions

OpenAI Chat Completions-compatible generation over messages[]. Read choices[].message.content for text; tool-capable responses can instead carry tool_calls. A content item can be text, image_url with a data URL, or input_video with base64 video bytes on macOS and Linux builds with native video support, which also requires ffmpeg and ffprobe on the service machine. Streaming uses chat.completion.chunk events terminated by data: [DONE]; do not apply that parser to the other protocols. Each endpoint requires a model whose supported_endpoints includes that exact path. These APIs forward the selected model's protocol, so optional parameters and response details can vary by local worker or remote provider. Local models keep inputs on the device; a configured remote model receives the request sent to it.

Guide and examplesLanguage and visionLocal Gateway APIAll endpoints

Request body

application/jsonRequired

modelstringRequiredText model ID from the model catalog. Its supported_endpoints must contain this path.
messagesobject[]RequiredConversation turns in order, oldest first.
rolestringRequiredMessage role, such as user or assistant.
contentstring | object[]RequiredMessage text, or an array of content items for multimodal input.
max_tokensintegerUpper bound on generated tokens. Omit it to use the model's own limit.
streambooleantrue returns Server-Sent Events instead of a single JSON body.

Responses

200The completion, or an SSE stream when stream: true.
idstring
objectstringchat.completion
modelstring
choicesobject[]
indexinteger
messageobject
finish_reasonstring | null
JSON
{
  "id": "chatcmpl_example",
  "object": "chat.completion",
  "model": "Qwen/Qwen3.5-4B",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello!"
      },
      "finish_reason": "stop"
    }
  ]
}
400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.
401Invalid or missing API key.
403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.
413Request body exceeds 512 MiB.
503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.