Realtime 实时会话
会话排队
CLI 与 Headless 服务在 --realtime-concurrency 满载时会排队等待,不立即返回 service_busy。排队发生在 HTTP 升级之前,此时尚未产生 session.created 事件。客户端应为 WebSocket 握手预留足够的超时预算,并在不需要时主动取消连接尝试。
连接建立后,按照下文示例收发事件。当前排队状态 runtime.queues.http_realtime 可通过运行状态查询查看。长时间空闲的应用应关闭不再使用的会话,释放会话名额和本机转录引擎资源。
完成一次文字会话
WS /v1/realtime
WebSocket /v1/realtime 使用事件协议,不是 JSON POST。先加载兼容的本机语言模型。将下列代码保存为 realtime-example.mjs,执行 export EDGESPEAK_MODEL="Qwen/Qwen3.5-4B"(替换为实际已加载的模型 ID),然后运行 node realtime-example.mjs。本示例为纯文字交互,需要 Node.js 22+,不需要麦克风或播报音色。
请求示例:
// Save as realtime-example.mjs; requires Node.js 22+.
const base = new URL(process.env.EDGESPEAK_BASE_URL || "http://127.0.0.1:1117/v1");
base.protocol = base.protocol === "https:" ? "wss:" : "ws:";
base.pathname = base.pathname.replace(/\/$/, "") + "/realtime";
if (!key) throw new Error("Set EDGESPEAK_API_KEY to the gateway API key");
const model = process.env.EDGESPEAK_MODEL;
if (!model) throw new Error("Set EDGESPEAK_MODEL to a loaded catalog model ID");
const ws = new WebSocket(base, ["realtime"]);
const send = (event) => ws.send(JSON.stringify(event));
let requested = false;
let finished = false;
const timer = setTimeout(() => {
console.error("Timed out before response.done");
process.exitCode = 1;
ws.close();
}, 120000);
ws.onerror = () => { console.error("WebSocket connection failed"); process.exitCode = 1; };
ws.onclose = () => {
clearTimeout(timer);
if (!finished) { console.error("Closed before successful completion"); process.exitCode = 1; }
};
ws.onmessage = ({ data }) => {
const event = JSON.parse(data);
if (event.type === "error") {
console.error(event.error); process.exitCode = 1; ws.close(); return;
}
if (event.type === "session.created") send({
type: "session.update",
session: { type: "realtime", model, output_modalities: ["text"] }
});
if (event.type === "session.updated") send({
type: "conversation.item.create",
item: { type: "message", role: "user",
content: [{ type: "input_text", text: "Say hello in one short sentence." }] }
});
if (event.type === "conversation.item.created" && !requested) {
requested = true;
send({ type: "response.create" });
}
if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
if (event.type === "response.done") {
finished = event.response?.status === "completed";
console.log("\nresponse.done:", event.response?.status);
if (!finished) process.exitCode = 1;
ws.close();
}
};输出示例:
Hello!
response.done: completed控制台输出仅为示意。事件交互顺序为 session.created → session.update → session.updated → conversation.item.create → conversation.item.created → response.create → response.output_text.delta → response.done,收到确认事件后再请求生成。开启鉴权后,服务端程序可使用普通鉴权请求头;浏览器 WebSocket 则使用「使用 API Key」模式展示的子协议传递 API Key。包含 API Key 的脚本仅用于可信环境。
发送音频进行转录
音频转录会话把上文示例的会话配置与消息发送分支换成下面的事件顺序。每行 JSON 均为一条独立的 WebSocket 消息;<base64 PCM16LE mono bytes> 需要替换为原始音频字节的 Base64 编码,而非 WAV 文件或文件路径。收到 session.updated 确认后再发送音频数据。
请求示例:
{"type":"session.update","session":{"type":"transcription","audio":{"input":{"format":{"type":"audio/pcm","rate":16000},"turn_detection":null}}}}
{"type":"input_audio_buffer.append","audio":"<base64 PCM16LE mono bytes>"}
{"type":"input_audio_buffer.commit"}输出示例:
{"type":"input_audio_buffer.committed","item_id":"item_example"}
{"type":"conversation.item.input_audio_transcription.completed","item_id":"item_example","transcript":"Hello world."}以上为事件节选,事件 ID 与转录文本因输入音频而异。转录会话仅输出文本,提交音频不会触发 conversation.item.created 事件。PCM16LE 单声道支持声明 8–192 kHz,默认为 16 kHz;G.711 使用 audio/pcmu 或 audio/pcma,固定为 8 kHz,直接省略 rate。可使用 ffmpeg 转换音频:ffmpeg -i hello.wav -f s16le -ac 1 -ar 16000 hello.pcm。将 hello.pcm 编码为 Base64 后分块追加并提交。需要固定说话语言时,把 session.audio.input.transcription.language 设为转录语言代码,例如 yue;省略则自动检测。
请求语音回复
转录 → 语言模型 → 播报链路需要设置 type:"realtime"、选择 model、设置 output_modalities:["audio"],并指定 audio.output.model 与兼容的 voice。audio.output.language 声明对话语言,写法同语言代码;在识别出说话语言之前,内置音色按它选择参考 profile。在手动模式下,追加音频并提交后需要发送 response.create,仅提交音频不会请求回复。若启用 server_vad,则会自动提交并回复,相关时长单位为秒。
请求示例:
{
"type": "session.update",
"session": {
"type": "realtime",
"model": "Qwen/Qwen3.5-4B",
"output_modalities": [
"audio"
],
"audio": {
"input": {
"format": {
"type": "audio/pcm",
"rate": 16000
},
"turn_detection": null
},
"output": {
"model": "Qwen/Qwen3-TTS-0.6B-Base",
"voice": "builtin:bright-girl"
}
}
}
}输出示例:
{"type":"response.output_audio.delta","delta":"<base64 PCM16LE mono bytes>","sample_rate":24000}
{"type":"response.done","response":{"status":"completed"}}以上输出仅为节选,并非完整事件 schema。解码音频时,请按 session.audio.output.format.rate 回显或 delta 中的 sample_rate 处理,不要硬编码为 24 kHz。Realtime delta 输出的是原始 PCM 数据,而 HTTP 播报 SSE 包含 WAV 文件头。所有会话配置应在发送首帧音频前完成;若后续修改会话,需要重新建立连接。直接向模型输入音频要求模型目录声明 audio_input,不支持原生视频。HTTP 事件流另见转录、播报、语言生成。
播报风格与克隆档位
session.audio.output 另有三个播报控制字段:
| 字段 | 说明 |
|---|---|
instructions | 字符串或 null。播报怎么说:交给播报模型的风格指令,与 /v1/audio/speech 的 instructions 同义。省略则不改;null 或空白表示回到音色保存的风格。只有 /v1/models 中 features 含 instruct 的播报模型接受非空值,否则当场报 errors.broadcast.instructUnsupported,整条更新不生效。它不是 session.instructions,后者是语言模型的系统提示,决定说什么。不在 session.created 中回显。 |
disable_style | 布尔。true 表示不用任何风格,包括音色保存的风格;false 表示回到音色保存的风格。不能在给了非空 instructions 的同时为 true。 |
clone_mode | "quick"、"ultimate" 或 null,缺省为 null(用该模型缺省档)。取值与 /v1/audio/speech 的 clone_mode 相同,各模型的档位见 /v1/models 的 clone_mode。点名了当前播报模型兑现不了的档位时返回 clone_mode_unsupported_by_model,message 以这个码开头并列出支持的档位,不会静默回落。 |
首帧音频之前可以多次发送 session.update,每条都是增量更新:省略字段保留旧值,显式传 null 才清空;音频开始后再发会返回 session_already_started,需要重新连接。更换 audio.output.model 而没有同时给 clone_mode 时,已设的 clone_mode 自动清空;换模型与点名档位可以放在同一条 session.update 里,按更换后的模型校验。
audio.output.voice 不接受 builtin:design,传入返回 invalid_session_field;instructions 不是字符串或 null、disable_style 不是布尔、disable_style: true 同时给了非空 instructions、传了已删除的 style_instruction、clone_mode 不是上述取值时,同样返回 invalid_session_field。按描述设计声音请使用 HTTP 播报。
建会话时的模型准备
session.update 依次准备识别模型、语言模型和播报模型,全部就绪后才回 session.updated。桌面 App 与 Headless 服务按同一张表报码:
| 原因 | 识别模型 | 语言模型 | 播报模型 |
|---|---|---|---|
| 点名的模型不存在(识别含云端模型) | model_not_found | realtime_language_model_unavailable | model_not_found |
| 模型没下载 / 没安装 | realtime_model_not_loaded | realtime_model_not_loaded | realtime_model_not_loaded |
| 模型正在后台下载 | — | model_downloading | — |
| 内存(或显存)不足,装不上 | model_load_insufficient_memory | language_model_insufficient_memory | model_load_insufficient_memory |
| 本机引擎被别的任务占着,等满时限仍没空 | service_busy | service_busy | — |
| 授权问题 | license_* | license_* | license_* |
| 其余装载失败(含模型文件损坏、磁盘空间不足) | realtime_runtime_unavailable | realtime_language_model_unavailable | realtime_tts_model_unavailable |
两个等待时限:
- 语言模型:准备总时限 110 秒,包括启动、偶发失败后的一次自动重试、等模型忙完别的请求,以及就绪后的预热。时限内没就绪返回
realtime_language_model_unavailable。首次加载大模型可能超时,建议建会话前先加载模型。 - 识别模型:本机引擎被占着时最多等 180 秒,等满仍没空返回
service_busy。
处理方式:model_not_found 换一个本机已安装的模型;realtime_model_not_loaded 先装好模型再重发 session.update,重试不会自愈;model_downloading 等下载完成后重发;service_busy 可以稍后重试;license_* 按 message 的指引处理。