Speech and voice library

List voices and choose a compatible pair

GET /v1/audio/voices

List voices before requesting synthesis. Pick a voice with available: true and inspect its per-model compatibility list; availability alone does not guarantee compatibility with every speech model. The following response excerpt shows one user voice; IDs are illustrative.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/voices"

Example output:

JSON
{
  "voices": [
    {
      "id": "user:00000000-0000-4000-8000-000000000001",
      "names": {
        "en-US": "My voice"
      },
      "descriptions": {},
      "supported_languages": [
        "en-US"
      ],
      "origin": "cloned",
      "compatibility": [],
      "available": true,
      "created_by_user": true
    }
  ]
}

Names and descriptions are indexed by locale, not a single name string. The empty compatibility list in this illustrative entry does not establish support for any model; provide an actually compatible model and voice pair in subsequent requests.

Generate speech as WAV

POST /v1/audio/speech

Required JSON request fields: model, voice, and input (or non-empty segments[]). language is optional; omit it or send auto to let the service choose, and see speech languages for what each model supports. The model and voice pair below is only an example; confirm that it exists and is compatible in the model catalog first. The response format is WAV.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3-TTS-0.6B-Base","voice":"builtin:bright-girl","input":"Hello world.","response_format":"wav"}' \
  -o speech.wav
file speech.wav

Example output:

speech.wav: RIFF (little-endian) data, WAVE audio, Microsoft PCM, 16 bit, mono 24000 Hz

This is example output from file, not JSON. The HTTP body is binary audio/wav data, streamed by default. Native sample rate depends on the model; omit sample_rate. The desktop gateway streaming interface only accepts an explicit value that matches the native sample rate, while the Headless service rejects explicit sample_rate. When a request fails, --fail-with-body exits with a non-zero status; inspect the saved error content rather than playing the file directly. See model-dependent speech controls.

Receive speech as SSE

When an event stream is needed, pass stream_format: "sse" and use curl -N to disable buffering. Audio data in each delta event is base64-encoded; the first decoded chunk includes the WAV header. Concatenate the decoded bytes in sequence to reconstruct the audio.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3-TTS-0.6B-Base","voice":"builtin:bright-girl","input":"Hello world.","stream_format":"sse","response_format":"wav"}' \
  -N

Example output:

event: speech.audio.delta
data: {"type":"speech.audio.delta","audio":"<base64 WAV header and PCM bytes>"}

event: speech.audio.done
data: {"type":"speech.audio.done","diagnostics":{},"usage":{"input_tokens":0,"output_tokens":0,"total_tokens":0}}

The event excerpt above omits internal diagnostics fields and uses placeholder audio. Successful synthesis ends with a speech.audio.done event; clients must also handle error events and unexpected disconnections before completion. Token counts in usage are zero here and cannot be used to estimate audio duration.

Style instructions, clone modes, and voice design

These fields go at the top level of the POST /v1/audio/speech JSON body. Each segments[] item can also carry its own instructions or disable_style; a segment that sets neither (including "", null, or false) inherits the top-level style.

ParameterDescription
instructionsThe OpenAI parameter of the same name, as a string. Omitted, "", or null keeps the style saved with the voice. Non-empty text uses that style for this request only. Only models whose /v1/models features include instruct accept non-empty text; other models return 400 instead of silently ignoring it. With voice: "builtin:design" it is the voice description and is required. On FireRedTeam/FireRedTTS3-Instruct and k2-fsa/OmniVoice the description only works with voice: "builtin:design": a specific voice (a preset voice, a saved voice, or builtin:auto) plus non-empty text returns 400 instead of silently swapping the selected voice; leave it empty to use the selected voice as usual.
disable_styleBoolean. true speaks this request without any style; false or omitted has no effect. Sending it together with a non-empty instructions returns 400 conflicting_speech_instructions.
clone_modeCloning mode: quick starts playback sooner, ultimate gives better similarity. One parameter shared by every cloning model. When omitted, the model's default mode is used. Naming a mode that cannot be delivered returns 400; it never silently falls back to the other mode.

/v1/models lists the cloning modes per model: speech models with modes carry clone_mode.values (available modes) and clone_mode.default_value (default mode); models without modes have no clone_mode field. Currently Qwen3-TTS Base (0.6B and 1.7B) and openbmb/VoxCPM2 offer both modes, with quick as the default.

JSON
{"id":"openbmb/VoxCPM2","clone_mode":{"values":["quick","ultimate"],"default_value":"quick"}}

The example is an excerpt of a catalog entry with other fields omitted.

Design a voice from a description

voice: "builtin:design" attaches no built-in or user voice; the description in instructions is the voice:

  • Supported models: openbmb/VoxCPM2, Qwen/Qwen3-TTS-1.7B-VoiceDesign, BreezeBlue/Breeze-TTS-2, FireRedTeam/FireRedTTS3-Instruct, and k2-fsa/OmniVoice. Other models reject the request instead of substituting a built-in voice.
  • A description is required, up to 120 characters; FireRedTeam/FireRedTTS3-Instruct allows up to 1000 and k2-fsa/OmniVoice up to 80.
  • FireRedTeam/FireRedTTS3-Instruct and k2-fsa/OmniVoice take a description only this way: a specific voice plus non-empty instructions returns 400; see the error table below.
  • clone_mode is not accepted.
  • Realtime sessions do not accept builtin:design; see Request spoken replies.

A k2-fsa/OmniVoice description is a list of attribute phrases separated by commas (full-width commas in Chinese), at most one per category: gender (male / female), age (child, teenager, young adult, middle-aged, elderly), pitch (very low pitch … very high pitch), whisper, an English accent (such as british accent, English only), and a Chinese dialect (such as 四川话, Chinese only). Voice design supports Chinese and English only; other languages may be unstable.

In the desktop app, filling in a description on the speech page switches the request to builtin:design automatically, and the selected voice is not used for that request.

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{"model":"FireRedTeam/FireRedTTS3-Instruct","voice":"builtin:design","instructions":"A calm, deep male narrator","input":"Hello, this is EdgeSpeak.","response_format":"wav"}' \
  -o designed.wav
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{"model":"k2-fsa/OmniVoice","voice":"builtin:design","instructions":"female, young adult, low pitch","input":"Hello, this is EdgeSpeak.","response_format":"wav"}' \
  -o omnivoice-designed.wav

Style and clone-mode errors

Errors use the standard envelope. Branch on error.code; do not match message text:

JSON
{"error":{"message":"...","type":"invalid_request_error","code":"conflicting_speech_instructions"}}
Caseerror.codeWhere the reason is
disable_style: true sent with a non-empty instructionsconflicting_speech_instructionsThe code itself
Request contains the removed style_instruction (top level or in segments[], any value)bad_requestmessage says to use instructions / disable_style
Non-empty instructions on a model that does not support style instructions (including FireRedTeam/FireRedTTS3-Instruct and k2-fsa/OmniVoice with a specific voice) / required description missing / description too longbad_requestStable key in message: errors.broadcast.instructUnsupported, errors.broadcast.instructRequired, or errors.broadcast.instructTooLong
Model lacks the named cloning mode / model has it but this voice does notbad_requestmessage starts with clone_mode_unsupported_by_model or clone_mode_unsupported_by_voice and lists the supported modes

These are problems with the request itself. Change the request; retrying it unchanged will not succeed.

Find model-specific speech examples

GET /v1/audio/speech/examples

Optional query parameters: model, mode, language, id, recipe. Query the full catalog first to get valid IDs and modes; passing invalid filters returns 400. The following command saves the full catalog first, then displays available modes using jq.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/speech/examples" \
  -o speech-examples.json
jq '.examples | map(.mode) | unique' speech-examples.json

Example output:

JSON
[
  "emotion-vector",
  "style-control",
  "voice-design"
]

This is example output processed by jq. The saved response contains schema_version, catalog_version, verified_at, resolved_language, models[], and examples[]. Each example includes applicable and reason; only use control options supported by the current model when making requests. For example, append ?model=voxcpm2&mode=voice-design&language=en-US to filter by criteria.

Create a reusable user voice

POST /v1/audio/voices

Saves a voice to the service host's library. Submit parameters using multipart format:

ParameterDescription
audio_sampleRequired, up to 10 MiB; use a short, clear reference clip.
ref_text / nameRequired, corresponding to the reference transcript and voice name, respectively.
consentRequired, must be true.
languageOptional, default zh-CN; must be set explicitly for other languages. Stored as a locale such as en-US; see Languages.
speaker_descriptionOptional speaker description.

Do not supply model.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/voices" \
  -F audio_sample=@reference.wav -F ref_text="Hello, this is my reference recording." \
  -F name="My voice" -F language=en-US -F consent=true \
  -o created-voice.json
cat created-voice.json

Example output:

JSON
{
  "success": true,
  "voice": {
    "id": "user:00000000-0000-4000-8000-000000000001",
    "names": {
      "en-US": "My voice"
    },
    "descriptions": {},
    "supported_languages": [
      "en-US"
    ],
    "origin": "cloned",
    "compatibility": [],
    "available": true,
    "created_by_user": true
  }
}

Use the actual returned voice.id for subsequent synthesis or deletion, not the illustrative UUID. Inspect the returned compatibility information before requesting synthesis.

Delete a user voice

DELETE /v1/audio/voices/:voice_id

Deletes the voice created above from the voice library. Run this only when you are certain it is no longer needed. Built-in voices cannot be deleted. The following example uses jq to read the saved response to retrieve the voice ID.

Request:

Shell
VOICE_ID=$(jq -er '.voice.id' created-voice.json)
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/voices/$VOICE_ID" \
  -X DELETE

Example output:

JSON
{
  "success": true,
  "deleted": {
    "id": "user:00000000-0000-4000-8000-000000000001",
    "name": "My voice"
  }
}

Desktop request queue

Local speech requests on the desktop share one queue. Clients submit concurrently, models execute in turn, and a waiting request does not hold a model. The model sleep setting controls idle eviction only; it does not keep several speech models loaded at once.

The request format is unchanged. If another speech request is ahead of this one, time to first audio includes queue time:

Shell
curl --no-buffer http://127.0.0.1:1117/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"openbmb/VoxCPM2","voice":"builtin:bright-girl","input":"你好。","response_format":"wav","stream_format":"sse"}'

When desktop gateway authentication is enabled, add Bearer credentials. Successful requests likewise send SSE events via speech.audio.delta and speech.audio.done; failures after audio transmission starts still use existing in-stream error handling.

Headless model queue

In the CLI and Headless service, REST speech, Realtime speech, MCP synthesis, voice preparation, and TTS preloading share one model execution queue. Requests execute in turn, and raising --speech-concurrency does not switch models while an earlier request is preparing or synthesizing. Cancelling a waiting request stops the server from loading the model; cancelling a running one waits for the native operation to finish before releasing model ownership. Native failures or cleanup timeouts can still return errors. Realtime handshakes also wait when at capacity; see the session queue.

Use the WAV or SSE request examples above; queueing adds no request fields, SSE events, or response fields. Client timeouts should cover queueing, model loading, and inference; see CLI concurrency options.

Generation timing and IndexTTS streaming

diagnostics.infer_seconds measures wall-clock time in the native synthesis stage, including vocoder computation and in-stage waits. It is neither pure GPU compute time nor model loading time, and it includes batch pre-generation time on non-streaming long-form text. Client time to first audio is measured from when the request is sent and may include queueing, loading, and voice preparation time.

The following request follows the existing protocol; replace input with your actual long text and select IndexTTS 2.5 from the model catalog. Streaming applies only when the text is split into multiple segments; a short sentence may return only a single audio event.

JSON
{"model":"IndexTeam/IndexTTS-2.5","voice":"builtin:bright-girl","input":"你好,这是一条语音速度测试。","language":"zh-CN","response_format":"wav","stream_format":"sse"}

Successful terminal-event diagnostic excerpt (illustrative values; other fields omitted):

JSON
{"type":"speech.audio.done","diagnostics":{"infer_seconds":6.0,"duration_seconds":3.0}}

The generation RTF in this example is 6.0 / 3.0 = 2.0; total RTF must be calculated separately using the client's full request duration. There is currently no public field for per-request model loading time, and total time minus generation time cannot be used to represent load time.

Continue: Language and vision · API index