Generate with Messages

POST /v1/messages

Anthropic Messages-compatible generation. Supply model, messages and max_tokens. This endpoint accepts x-api-key as an alternative to Bearer authentication and forwards anthropic-version and anthropic-beta when present. Read text blocks inside content[]; this shape differs from both Responses and Chat Completions, and errors use the Anthropic error envelope. Streaming uses content_block_delta and message_stop. Each endpoint requires a model whose supported_endpoints includes that exact path. These APIs forward the selected model's protocol, so optional parameters and response details can vary by local worker or remote provider. Local models keep inputs on the device; a configured remote model receives the request sent to it.

Guide and examplesLanguage and visionAll endpoints

Parameters

anthropic-versionstringForwarded to the model when present, for example 2023-06-01.
anthropic-betastringForwarded to the model when present.

Request body

application/jsonRequired

modelstringRequiredText model ID from the model catalog. Its supported_endpoints must contain this path.
messagesobject[]RequiredConversation turns in order, oldest first.
rolestringRequiredMessage role, such as user or assistant.
contentstring | object[]RequiredMessage text, or an array of content items for multimodal input.
max_tokensintegerRequiredRequired upper bound on generated tokens.
streambooleantrue returns Server-Sent Events instead of a single JSON body.

Responses

200The message, or an SSE stream when stream: true.
idstring
typestringmessage
rolestring
modelstring
contentobject[]
typestring
textstring
stop_reasonstring | null
usageobject
input_tokensinteger
output_tokensinteger
JSON
{
  "id": "msg_example",
  "type": "message",
  "role": "assistant",
  "model": "Qwen/Qwen3.5-4B",
  "content": [
    {
      "type": "text",
      "text": "Hello!"
    }
  ],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 16,
    "output_tokens": 3
  }
}
400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.
401Invalid or missing API key.
403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.
413Request body exceeds 512 MiB.
503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.