Realtime sessions
Session queue
CLI and Headless services queue when --realtime-concurrency is full instead of immediately returning service_busy. Queuing happens before the HTTP upgrade, so no session.created event exists yet. Clients should allow sufficient timeout budget for WebSocket handshakes and cancel connection attempts that are no longer needed.
Once connected, exchange events as shown in the example below. Inspect the current queue status in runtime.queues.http_realtime through a running-state request. Applications that sit idle should close unused sessions to release session capacity and local transcription engine resources.
Run one complete text exchange
WS /v1/realtime
WebSocket /v1/realtime uses events rather than a JSON POST. First load a compatible local language model. Save the following code as realtime-example.mjs; run export EDGESPEAK_MODEL="Qwen/Qwen3.5-4B" with your actual loaded model ID, then run node realtime-example.mjs. This text-only example requires Node.js 22+ and no microphone or speech voice.
Request:
// Save as realtime-example.mjs; requires Node.js 22+.
const base = new URL(process.env.EDGESPEAK_BASE_URL || "http://127.0.0.1:1117/v1");
base.protocol = base.protocol === "https:" ? "wss:" : "ws:";
base.pathname = base.pathname.replace(/\/$/, "") + "/realtime";
if (!key) throw new Error("Set EDGESPEAK_API_KEY to the gateway API key");
const model = process.env.EDGESPEAK_MODEL;
if (!model) throw new Error("Set EDGESPEAK_MODEL to a loaded catalog model ID");
const ws = new WebSocket(base, ["realtime"]);
const send = (event) => ws.send(JSON.stringify(event));
let requested = false;
let finished = false;
const timer = setTimeout(() => {
console.error("Timed out before response.done");
process.exitCode = 1;
ws.close();
}, 120000);
ws.onerror = () => { console.error("WebSocket connection failed"); process.exitCode = 1; };
ws.onclose = () => {
clearTimeout(timer);
if (!finished) { console.error("Closed before successful completion"); process.exitCode = 1; }
};
ws.onmessage = ({ data }) => {
const event = JSON.parse(data);
if (event.type === "error") {
console.error(event.error); process.exitCode = 1; ws.close(); return;
}
if (event.type === "session.created") send({
type: "session.update",
session: { type: "realtime", model, output_modalities: ["text"] }
});
if (event.type === "session.updated") send({
type: "conversation.item.create",
item: { type: "message", role: "user",
content: [{ type: "input_text", text: "Say hello in one short sentence." }] }
});
if (event.type === "conversation.item.created" && !requested) {
requested = true;
send({ type: "response.create" });
}
if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
if (event.type === "response.done") {
finished = event.response?.status === "completed";
console.log("\nresponse.done:", event.response?.status);
if (!finished) process.exitCode = 1;
ws.close();
}
};Example output:
Hello!
response.done: completedThe console output is illustrative. The event interaction sequence is session.created → session.update → session.updated → conversation.item.create → conversation.item.created → response.create → response.output_text.delta → response.done; wait for confirmation events before requesting generation. With authentication enabled, server applications can use normal auth headers; browser WebSocket clients use the subprotocol shown in “Use API Key” mode to pass the API key. Keep scripts containing API keys in trusted environments.
Send audio for transcription
For an audio transcription session, swap the session configuration and message-sending branches in the example above for the event sequence below. Each JSON line is a separate WebSocket message; <base64 PCM16LE mono bytes> must be replaced with the base64 encoding of raw audio bytes, not a WAV file or file path. Wait for the session.updated confirmation before sending audio data.
Request:
{"type":"session.update","session":{"type":"transcription","audio":{"input":{"format":{"type":"audio/pcm","rate":16000},"turn_detection":null}}}}
{"type":"input_audio_buffer.append","audio":"<base64 PCM16LE mono bytes>"}
{"type":"input_audio_buffer.commit"}Example output:
{"type":"input_audio_buffer.committed","item_id":"item_example"}
{"type":"conversation.item.input_audio_transcription.completed","item_id":"item_example","transcript":"Hello world."}These are event excerpts; event IDs and transcripts vary by input audio. A transcription session returns text only; committing audio does not trigger a conversation.item.created event. PCM16LE mono supports declared rates of 8–192 kHz, default 16 kHz; G.711 uses audio/pcmu or audio/pcma fixed at 8 kHz, omitting rate. Convert audio with ffmpeg: ffmpeg -i hello.wav -f s16le -ac 1 -ar 16000 hello.pcm. base64-encode hello.pcm, append it in chunks, and commit. To pin the spoken language, set session.audio.input.transcription.language to a transcription language code such as yue; omit it to detect the language.
Request a spoken reply
For transcription → language model → speech, set type:"realtime", select model, set output_modalities:["audio"], and specify audio.output.model with a compatible voice. audio.output.language declares the conversation language in the shared code format; built-in voices use it to choose a reference profile until a spoken language has been detected. In manual mode, append audio and commit, then send response.create; committing audio alone does not request a reply. If server_vad is enabled, it automatically commits and responds; durations are in seconds.
Request:
{
"type": "session.update",
"session": {
"type": "realtime",
"model": "Qwen/Qwen3.5-4B",
"output_modalities": [
"audio"
],
"audio": {
"input": {
"format": {
"type": "audio/pcm",
"rate": 16000
},
"turn_detection": null
},
"output": {
"model": "Qwen/Qwen3-TTS-0.6B-Base",
"voice": "builtin:bright-girl"
}
}
}
}Example output:
{"type":"response.output_audio.delta","delta":"<base64 PCM16LE mono bytes>","sample_rate":24000}
{"type":"response.done","response":{"status":"completed"}}This output is an excerpt, not a complete event schema. When decoding audio, use the rate echoed in session.audio.output.format.rate or the sample_rate in the delta; do not hard-code 24 kHz. Realtime deltas output raw PCM data, whereas HTTP speech SSE includes a WAV header. Complete all session configuration before sending the first audio frame; modifying the session later requires reconnecting. Direct model audio input requires the model catalog to declare audio_input; native video is unsupported. For HTTP event streams see transcription, speech and language generation.
Speech style and clone mode
session.audio.output has three more speech controls:
| Field | Description |
|---|---|
instructions | String or null. How the speech sounds: a speaking-style instruction for the speech model, same meaning as instructions on /v1/audio/speech. Omitted: no change. null or blank: back to the style saved with the voice. Only speech models whose /v1/models features include instruct accept a non-empty value; otherwise the update is rejected right away with errors.broadcast.instructUnsupported and does not apply. Not the same as session.instructions, which is the language model's system prompt (what to say). Not echoed in session.created. |
disable_style | Boolean. true: speak without any style, including the one saved with the voice. false: back to the voice's saved style. Cannot be true together with a non-empty instructions. |
clone_mode | "quick", "ultimate", or null; defaults to null (the model's default mode). Same values as clone_mode on /v1/audio/speech; see clone_mode in /v1/models for each model's modes. Naming a mode the current speech model cannot deliver returns clone_mode_unsupported_by_model, whose message starts with that code and lists the supported modes; it never silently falls back. |
Before the first audio frame you can send session.update several times, and each one is an incremental update: omitted fields keep their previous values, and only an explicit null clears one. Once audio has started, another session.update returns session_already_started and you need to reconnect. Changing audio.output.model without also sending clone_mode clears any clone_mode set earlier. You can change the model and name a mode in the same session.update; the mode is validated against the new model.
audio.output.voice does not accept builtin:design and returns invalid_session_field. An instructions that is not a string or null, a disable_style that is not a boolean, disable_style: true with a non-empty instructions, the removed style_instruction field, or a clone_mode outside the values above also returns invalid_session_field. To design a voice from a description, use HTTP speech.
Model preparation during session setup
session.update prepares the transcription, language, and speech models in turn and replies with session.updated only when all are ready. The desktop app and the Headless service report failures with the same table:
| Cause | Transcription model | Language model | Speech model |
|---|---|---|---|
| Named model does not exist (transcription includes cloud models) | model_not_found | realtime_language_model_unavailable | model_not_found |
| Model not downloaded / not installed | realtime_model_not_loaded | realtime_model_not_loaded | realtime_model_not_loaded |
| Model is downloading in the background | — | model_downloading | — |
| Not enough memory (or VRAM) to load | model_load_insufficient_memory | language_model_insufficient_memory | model_load_insufficient_memory |
| Local engine busy with other work past the wait limit | service_busy | service_busy | — |
| License problem | license_* | license_* | license_* |
| Any other load failure (including corrupt model files or low disk space) | realtime_runtime_unavailable | realtime_language_model_unavailable | realtime_tts_model_unavailable |
Two wait limits apply:
- Language model: 110 seconds in total, covering startup, one automatic retry after an intermittent failure, waiting for the model to finish other requests, and warm-up once ready. If it is not ready in time, the result is
realtime_language_model_unavailable. A first load of a large model can exceed this, so load the model before opening the session. - Transcription model: waits up to 180 seconds while the local engine is busy, then returns
service_busy.
How to respond: for model_not_found, pick a locally installed model; for realtime_model_not_loaded, install the model and resend session.update (retrying alone will not fix it); for model_downloading, resend after the download finishes; service_busy can be retried later; for license_*, follow the guidance in message.
Continue: Models and service · API index