Language and vision

Choose a compatible model first

Before calling, confirm model availability through model discovery and loading. The selected model's supported_endpoints must include the target endpoint path. These APIs directly forward the selected model’s protocol; optional parameters and response details can vary by local runtime or remote provider. The examples below show typical non-streaming response excerpts; actual generated text, token counts, or IDs may differ.

Generate with Responses

POST /v1/responses

Pass model and input; input can be a string or supported structured input items. Read generated text from message content inside output[], not a universal top-level text field.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/responses" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B","input":"Say hello in one short sentence.","max_output_tokens":128}'

Example output:

JSON
{
  "id": "resp_example",
  "object": "response",
  "status": "completed",
  "model": "Qwen/Qwen3.5-4B",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "status": "completed",
      "content": [
        {
          "type": "output_text",
          "text": "Hello!"
        }
      ]
    }
  ]
}

Reasoning-capable models may add reasoning items before the message. When parsing, inspect item types rather than assuming output[0] is text.

Generate with Chat Completions

POST /v1/chat/completions

Pass model and messages[] with roles and content, suitable for clients that manage chat history themselves.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B","messages":[{"role":"user","content":"Say hello in one short sentence."}],"max_tokens":128}'

Example output:

JSON
{
  "id": "chatcmpl_example",
  "object": "chat.completion",
  "model": "Qwen/Qwen3.5-4B",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello!"
      },
      "finish_reason": "stop"
    }
  ]
}

Read generated text from choices[].message.content. If the model called tools, the response also returns tool_calls.

Set the translation direction

POST /v1/responses and POST /v1/chat/completions accept an optional top-level translation object that sets the translation direction and style:

FieldDescription
target_languageLanguage to translate into for this turn, as a BCP-47 tag such as en, ja, or zh-TW.
language_pairAn array of two different languages, such as ["zh", "ja"]. When the source text is in one of them, it is translated into the other, which suits back-and-forth interpreting.
source_languageHint for the source language. With language_pair, it decides which item the source text is in; when omitted, this is inferred from the text.
styleTranslation style: natural, conversational, or faithful.

The target is chosen in this order: target_language first; then the item of language_pair that is not the source language, or the second item when the source cannot be matched to either; without both, Chinese source text (or a Chinese source_language) is translated into English and anything else into Chinese.

  • Local translation models: the translation request is built from these rules.
  • General chat models and cloud services are covered too: the language pair, target language, and style are written into the interpreting prompt sent to the model. For Chat they are appended to the first system message (one is added if missing); for Responses they are appended to instructions. The translation field itself is not forwarded to cloud services.
  • When translation is omitted or null, the request is left unchanged.
  • A malformed object (unknown field, wrong type, a language pair with the same language twice, or a style outside the values above) returns 400 translation_options_invalid; an unsupported language returns 400 translation_language_unsupported.
JSON
{
  "model": "Qwen/Qwen3.5-4B",
  "input": "我们下周一开会。",
  "translation": {
    "language_pair": ["zh", "ja"],
    "style": "conversational"
  }
}

This example translates the Chinese source into Japanese; Japanese source text would be translated into Chinese. The MCP tool edgespeak_create_response and the Realtime session.translation use the same shape.

Generate with Messages

POST /v1/messages

Use this for Anthropic-compatible clients. Pass model, messages, and max_tokens. The gateway supports x-api-key authentication and forwards anthropic-version and anthropic-beta when explicitly passed in the request. Before calling, confirm in the model catalog that the model supports this endpoint.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/messages" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B","max_tokens":128,"messages":[{"role":"user","content":"Say hello in one short sentence."}]}' \
  -H "anthropic-version: 2023-06-01"

Example output:

JSON
{
  "id": "msg_example",
  "type": "message",
  "role": "assistant",
  "model": "Qwen/Qwen3.5-4B",
  "content": [
    {
      "type": "text",
      "text": "Hello!"
    }
  ],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 16,
    "output_tokens": 3
  }
}

Read text blocks inside content[]. This response structure differs from both Responses and Chat Completions.

Tokenize text

POST /v1/tokenize

This example uses the local runtime’s content input. Token IDs depend on the selected model’s tokenizer. Remote providers may use a different tokenization format; the gateway forwards their supported format.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/tokenize" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B","content":"Hello world."}'

Example output:

JSON
{
  "tokens": [
    9707,
    1879,
    13
  ]
}

The returned array is illustrative; do not hard-code these IDs. If only an array is returned, you can count its length directly to get the token count. Tokenizing raw text does not include the additional tokens a chat template may insert.

Read language-model streams

Add stream:true and curl -N to generation calls. Parse SSE event boundaries, not individual network chunks. This example uses Chat Completions; other endpoints use their respective event names.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3.5-4B","messages":[{"role":"user","content":"Say hello."}],"stream":true,"max_tokens":128}' \
  -N

Example output:

data: {"id":"chatcmpl_example","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello!"},"finish_reason":null}]}

data: {"id":"chatcmpl_example","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

The above is an event stream excerpt. Event names differ across protocols: Responses typically uses response.output_text.delta and response.completed; Messages uses content_block_delta and message_stop. Do not apply Chat Completions’ [DONE] parsing logic to every protocol. Errors after streaming starts can also appear as a broken connection; check each protocol's own signals to confirm completion.

Send an image or video

Before calling, select a model advertising vision support. The following Python 3 script converts a local JPEG into a complete JSON request, avoiding OS-specific base64 command-line differences. Run it in the directory containing photo.jpg; the script generates vision-request.json.

Request:

Shell
python3 - <<'PYIMAGE'
import base64, json
from pathlib import Path
data = base64.b64encode(Path("photo.jpg").read_bytes()).decode("ascii")
body = {"model": "Qwen/Qwen3.5-4B", "messages": [{"role": "user", "content": [
    {"type": "text", "text": "Describe this image."},
    {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64," + data}}
]}], "max_tokens": 256}
Path("vision-request.json").write_text(json.dumps(body), encoding="utf-8")
PYIMAGE
curl --fail-with-body "$EDGESPEAK_BASE_URL/chat/completions" \
  -H "Content-Type: application/json" -d @vision-request.json

Example output:

JSON
{
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "A mountain lake under a clear sky."
      },
      "finish_reason": "stop"
    }
  ]
}

The output is a response excerpt; the description depends on the input image. Chat Completions supports passing native video directly, formatted as {"type":"input_video","input_video":{"data":"base64 video bytes"}}. For requirements, see video prerequisites, sampling and size limits and context window configuration. The Realtime endpoint does not accept native video.

Continue: Realtime sessions · API index