Generate speech as WAV

POST /v1/audio/speech

Synthesizes speech from text. model, voice and either input or a non-empty segments[] are required; obtain voice IDs from GET /v1/audio/voices. The default stream_format: "audio" streams binary audio/wav, and sse returns speech.audio.delta events followed by speech.audio.done, with base64 audio in each delta and the WAV header in the first one. Omit sample_rate for native-rate output: the desktop gateway accepts an explicit native rate only and returns 400 for streaming resampling, and headless hosts reject explicit values. Use only the controls the selected model supports.

Guide and examplesSpeech and voice libraryLocal Gateway APIAll endpoints

Request body

application/jsonRequired

modelstringRequiredSpeech model ID from the model catalog. One model serves the whole request, including every segment.
voicestringRequiredVoice ID from GET /v1/audio/voices, such as builtin:bright-girl or a saved user:<uuid>. builtin:auto lets the model design a voice, which Qwen/Qwen3-TTS-VoiceDesign requires; CustomVoice models fall back to their first official voice, Vivian.
inputstringSupply input or a non-empty segments[].
segmentsobject[]Multi-segment input, at least one entry. Each segment inherits the top-level voice and generation options unless it overrides them.
inputstringText for this segment.
voicestringVoice for this segment. Omit it to inherit the top-level voice.
instructionsstring | nullSpeaking style for this segment. Non-empty text overrides the top-level style; omitted, "", or null inherits it.
disable_stylebooleantrue speaks this segment without any style; false or omitted inherits the top-level style. Cannot be combined with a non-empty instructions.
speednumberPlayback rate for this segment. Omit it to inherit the top-level speed.
languagestringLanguage for this segment. Omit it to inherit the top-level language.
response_formatstringWAV is the supported response format.wav
stream_formatstringaudio streams the WAV bytes back directly; sse returns an event stream whose deltas carry base64 audio. Defaults to audio.audio | sseDefault "audio"
languagestringLanguage of the input text, such as zh or en-US. Omit it or send auto to choose automatically. Models without a language input still check it against their supported_languages and use it to pick reference audio. A language the model does not support returns 422 language_unsupported. See /docs/languages#speech for supported languages.
seedintegerSampling seed. Send the same seed to reproduce a result; omit it for a fresh result each time.
sample_rateintegerOmit for native-rate output.
guidance_scalenumberClassifier-free guidance strength. Higher values follow the voice and text more closely at the cost of naturalness. Models without a diffusion stage ignore it.
inference_stepsintegerDiffusion steps. More steps trade speed for quality. Models without a diffusion stage ignore it.
temperaturenumberSampling temperature, range 0.1-2. A value outside the range returns 400; models that do not sample ignore it.
top_pnumberNucleus sampling cutoff, range 0.1-1. A value outside the range returns 400; models that do not sample ignore it.
top_kintegerNumber of sampling candidates, range 1-100. A value outside the range returns 400; models that do not sample ignore it.
repetition_penaltynumberPenalty on repeated tokens, starting at 1 for no penalty. The upper bound depends on the model and a value outside the range returns 400; models that do not sample ignore it.
retry_badcasebooleanLet the model retry a generation it judges failed. openbmb/VoxCPM2 only; other models ignore it.
clone_recipestringopenbmb/VoxCPM2 clone recipe, controllable or ultimate. Omit it to use the voice's saved preference. ultimate has no style control, so pair a non-empty instructions with controllable. Another model returns 400 clone_recipe_requires_voxcpm2, and builtin:auto returns 400 clone_recipe_requires_saved_voice.
instructionsstring | nullNatural-language speaking style under the OpenAI field name, such as a tone or an emotion. Omitted, "", or null keeps the style saved with the voice; non-empty text uses that style for this request. With voice: "builtin:design" it is the required voice description. A model that does not support instructions returns 400 for non-empty text instead of dropping it. On FireRedTeam/FireRedTTS3-Instruct and k2-fsa/OmniVoice a description works only with voice: "builtin:design"; a specific voice plus non-empty text returns 400 bad_request (errors.broadcast.instructUnsupported). The removed style_instruction field returns 400 bad_request.
disable_stylebooleantrue speaks this request without any style; false or omitted has no effect. Sending true with a non-empty instructions returns 400 conflicting_speech_instructions.
speednumberPlayback rate, where 1 is the model's natural speed. Models without a speed control ignore it.
user_marksPronunciation and control input; shape depends on the model.
phonetic_spansPronunciation and control input; shape depends on the model.
control_policyPronunciation and control input; shape depends on the model.

Responses

200Generated audio. The default body is binary audio/wav; with stream_format: "sse" it is an event stream whose terminal event is speech.audio.done. Also handle an error event or a disconnection before done.
400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.
401Invalid or missing API key.
403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.
413Request body exceeds 512 MiB.
422The selected model does not cover this language (language_unsupported). Change the model, not the language code.
503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.