Languages
Use one set of short language codes across the APIs.
zh, yue, en or ja for transcription, alignment, text normalization, speech, Realtime, MCP tools and CLI options. Which other spellings each API accepts, and how it normalizes them, is described below; the speech example filter and the voice library are exceptions.This page covers the language of your audio or text. It is unrelated to the interface language of the app or the CLI, and to language models.
Code format
These rules apply to transcription, alignment, text normalization, speech synthesis and Realtime, and to the MCP tools and CLI options that call them. See Exceptions for the two places that do not normalize.
- Recommended form: the two-letter ISO 639-1 code. The only three-letter code is Cantonese,
yue. - Case and separators: matching ignores case and surrounding whitespace, and
_works like-.ZH,en_USandjaare all valid. - Region and script subtags: full BCP-47 tags are accepted and resolve to their language:
en-USanden-GB→en,pt-BR→pt,zh-CN,zh-Hans,zh-TWandzh-Hant-TW→zh. - Other spellings: ISO 639-3 codes (
eng,deu,cmn,fil) and English names (English,Mandarin,Cantonese) are also recognized. Prefer the short code in new integrations. - Responses: the
languagefield returned by transcription, alignment and text normalization is the resolved code, for examplezhoryue.
| You send | Resolves to |
|---|---|
zh, zh-CN, zh-Hans, zh-TW, zh-HK, zh-Hant, cmn, Mandarin | zh |
yue, zh-yue, Cantonese | yue |
en, en-US, en-GB, eng, English | en |
pt, pt-BR, pt-PT | pt |
tl, fil, Filipino | tl |
nb, nob, nor, Norwegian | nb |
zh-HK and zh-TW mean Chinese (Mandarin), not Cantonese. For Cantonese audio or text, send yue.Chinese script is not a separate language. Transcription and alignment treat zh-Hant as zh. In text normalization, an explicit Chinese tag decides the script of the readings it generates in tn mode: zh-Hant, or a TW, HK or MO region, gives Traditional characters. A script subtag wins over the region, so zh-Hans-TW stays Simplified. Text kept from the input is never converted, and itn output is not affected. Translation keeps Traditional Chinese as its own target, zh-hant.
Omitting it, und and auto
The code format is shared. What happens when you leave language out, and which special values exist, depends on the capability:
| Capability | Omitted | und | auto |
|---|---|---|---|
| Transcription | Detects the language from the audio | Rejected | Rejected |
| Alignment | Infers the language from the reference text; if it cannot, aligns without a language | No language information: each word is matched against every language, and normalization reads segment by segment | Rejected |
| Text normalization | Rejected, language is required | No language information: readings are chosen segment by segment from the text | Rejected |
| Speech | Chosen automatically | Not a speech value | Same as omitting it |
und is the BCP-47 tag for "undetermined". UND and und-Latn are treated as und.
Parameters by API
| API | Parameter | Required | When omitted | Unsupported value |
|---|---|---|---|---|
POST /v1/audio/transcriptions | language (multipart) | No | Detected from audio | 400 asr_language_unsupported, param=language |
POST /v1/audio/alignments | language (multipart) | No | Inferred from the reference text | 400 bad_request for a malformed tag; 422 language_unsupported when the selected model does not cover the language; both with param=language |
POST /v1/text/normalizations | language (JSON) | Yes | 400 normalization_invalid_request | 400 normalization_invalid_request for an unrecognized tag; 422 normalization_language_unsupported for a language without normalization rules |
POST /v1/audio/speech | language, segments[].language | No | Chosen automatically | 422 language_unsupported |
POST /v1/audio/voices | language (multipart) | No | Saved as zh-CN | — |
GET /v1/audio/speech/examples | language (query) | No | No filter | — |
WS /v1/realtime | session.audio.input.transcription.language | No | Detected from audio | — |
WS /v1/realtime | session.audio.output.language | No | Uses the detected spoken language | — |
In Realtime, session.audio.output.language is the conversation language: built-in voices use it to choose a reference profile until a spoken language has been detected, after which the detected language takes precedence. /v1/text/segmentations and the speaker endpoints do not take a language. Voice library entries store the tag as sent; see Exceptions.
MCP tools and CLI options pass the same values to the same capabilities:
| MCP tool | CLI option | Capability |
|---|---|---|
edgespeak_transcribe, edgespeak_transcribe_file: language | edgespeak-cli transcribe --language | Transcription |
edgespeak_align: language | edgespeak-cli align --language | Alignment |
edgespeak_create_speech: language | edgespeak-cli speech --language | Speech |
edgespeak_add_voice: language | edgespeak-cli voices add --language | Voice library, default zh-CN |
edgespeak_speech_examples: language | edgespeak-cli speech-examples --language | Example filter |
edgespeak-cli transcribe --language needs a reachable gateway, either the desktop app or edgespeak-cli serve. Without one, the command exits with an error instead of ignoring the option.
Exceptions
- Speech example filter (
GET /v1/audio/speech/examples,edgespeak_speech_examples,edgespeak-cli speech-examples): only checks whether the first subtag iszh. Every other value, includingcmn,Mandarin,yueandzhwith surrounding spaces, returns English examples. Sendzhoren. - Voice library (
POST /v1/audio/voices,edgespeak_add_voice,edgespeak-cli voices add): trims surrounding whitespace and saves the tag as sent. It becomes the voice'slocaleand the key ofnames;Englishis not converted toenoren-US. Send a locale such aszh-CNoren-US.
Common languages
With the default alignment model, alignment covers every language in this table.
| Language | Code | Also accepted | Transcription | Normalization | Translation |
|---|---|---|---|---|---|
| Chinese (Mandarin) | zh | zh-CN, zh-TW, zh-Hans, zh-Hant, cmn | ✓ | ✓ | ✓ |
| Cantonese | yue | zh-yue | ✓ | — | ✓ |
| English | en | en-US, en-GB | ✓ | ✓ | ✓ |
| Japanese | ja | ja-JP | ✓ | ✓ | ✓ |
| Korean | ko | ko-KR | ✓ | ✓ | ✓ |
| French | fr | fr-FR, fr-CA | ✓ | ✓ | ✓ |
| German | de | de-DE | ✓ | ✓ | ✓ |
| Spanish | es | es-ES, es-MX | ✓ | ✓ | ✓ |
| Portuguese | pt | pt-BR, pt-PT | ✓ | ✓ | ✓ |
| Russian | ru | ru-RU | ✓ | ✓ | ✓ |
| Arabic | ar | ar-SA | ✓ | ✓ | ✓ |
| Italian | it | it-IT | ✓ | ✓ | ✓ |
| Vietnamese | vi | vi-VN | ✓ | ✓ | ✓ |
| Thai | th | th-TH | ✓ | ✓ | ✓ |
| Indonesian | id | id-ID | ✓ | ✓ | ✓ |
| Hindi | hi | hi-IN | ✓ | ✓ | ✓ |
Speech support depends on the model; see Speech.
Supported languages by capability
The capabilities cover different sets. Check the one you call rather than assuming every language works everywhere.
Transcription
30 languages: zh en yue ar de fr es pt id it ko ru th vi ja tr hi ms nl sv da fi pl cs tl fa el hu mk ro.
The same list applies to every transcription model.
Alignment
| Model | Languages |
|---|---|
EdgeSpeak/Lattice-2 (default) | Every language in the catalog, mixed-language text, und, and other well-formed tags such as nan or wuu, which have no dictionary pronunciations and perform worse than und |
EdgeSpeak/Lattice-1 | zh, en, de only. und and tags outside the catalog return 400 bad_request; other catalog languages return 422 language_unsupported |
Tags outside the catalog keep their script and region subtags, for example nan-Hant. For how alignment applies text normalization, see Forced alignment.
Text normalization
55 languages, the same for tn and itn: ar az bg bs ca cs da de el en es et fa fi fr he hi hr hu hy id is it ja ka kk km ko lo lt lv mk mn mr ms my nb nl pl pt ro ru sk sl sq sr sv sw ta th tl tr uk vi zh.
Cantonese yue has no normalization rules. The running list and the categories per language are in the EdgeSpeak/Skylark entry of GET /v1/models: x_edgespeak.normalization.languages_by_mode and classes_by_language.
Speech
Each speech model supports its own set. Read supported_languages for the model in GET /v1/models; when x_edgespeak.speech_language.open_set is true, the model also accepts ISO 639 codes outside that list. Region and script variants of a listed language, such as zh-TW or en-US, are accepted, and auto always is.
| Model | Languages |
|---|---|
k2-fsa/OmniVoice | Open set: every catalog language plus other ISO 639 codes |
Qwen/Qwen3-TTS-* | zh en de fr es pt it ko ru ja |
IndexTeam/IndexTTS-2.5 | zh en ja es ar |
FireRedTeam/FireRedTTS3 | zh yue en de fr es it pt ru ja ko ar cs nl fi el hi id pl ro th tr uk vi, plus 21 Chinese dialects written zh-x-<name>, such as zh-x-sichuan |
FireRedTeam/FireRedTTS3-Instruct | Same 24 languages, without the dialects |
openbmb/VoxCPM2 | zh en de ja |
BreezeBlue/Breeze-TTS-2 | zh en |
FireRedTTS3-Instruct, VoxCPM2 and Breeze-TTS-2 do not take a language input; for them language selects the matching reference audio and is still checked against the list. List every model's languages from a running service:
curl --fail-with-body "$EDGESPEAK_BASE_URL/models" \
| jq '.data[] | select(.supported_languages) | {id, supported_languages}'Translation
38 target languages: the 36 marked in the catalog, plus Traditional Chinese zh-hant and Uyghur ug. Translation tags are stored in lowercase, as in the target_langs field of JSON exports. See Translation requirements for how translation works.
Language catalog
EdgeSpeak recognizes 82 languages. Any code, ISO 639-3 code or English name in this table is accepted as input; the default alignment model covers all of them. The last three columns show which languages the other capabilities support.
| Code | Language | ISO 639-3 | Transcription | Normalization | Translation |
|---|---|---|---|---|---|
af | Afrikaans | afr | |||
ar | Arabic | ara arb | ✓ | ✓ | ✓ |
az | Azerbaijani | aze azj | ✓ | ||
be | Belarusian | bel | |||
bg | Bulgarian | bul | ✓ | ||
bn | Bengali (Bangla) | ben | ✓ | ||
bo | Tibetan | bod tib | ✓ | ||
bs | Bosnian | bos | ✓ | ||
ca | Catalan | cat | ✓ | ||
cs | Czech | ces cze | ✓ | ✓ | ✓ |
cy | Welsh | cym wel | |||
da | Danish | dan | ✓ | ✓ | |
de | German | deu ger | ✓ | ✓ | ✓ |
el | Greek | ell gre | ✓ | ✓ | |
en | English | eng | ✓ | ✓ | ✓ |
es | Spanish (Castilian) | spa | ✓ | ✓ | ✓ |
et | Estonian | est ekk | ✓ | ||
eu | Basque | eus | |||
fa | Persian (Farsi) | fas pes | ✓ | ✓ | ✓ |
fi | Finnish | fin | ✓ | ✓ | |
fr | French | fra fre | ✓ | ✓ | ✓ |
ga | Irish | gle | |||
gu | Gujarati | guj | ✓ | ||
he | Hebrew | heb | ✓ | ✓ | |
hi | Hindi | hin | ✓ | ✓ | ✓ |
hr | Croatian | hrv | ✓ | ||
hu | Hungarian | hun | ✓ | ✓ | |
hy | Armenian | hye | ✓ | ||
id | Indonesian | ind | ✓ | ✓ | ✓ |
is | Icelandic | isl ice | ✓ | ||
it | Italian | ita | ✓ | ✓ | ✓ |
ja | Japanese | jpn | ✓ | ✓ | ✓ |
ka | Georgian | kat geo | ✓ | ||
kk | Kazakh | kaz | ✓ | ✓ | |
km | Khmer (Cambodian) | khm | ✓ | ✓ | |
kn | Kannada | kan | |||
ko | Korean | kor | ✓ | ✓ | ✓ |
lg | Ganda (Luganda) | lug | |||
lo | Lao | lao | ✓ | ||
lt | Lithuanian | lit | ✓ | ||
lv | Latvian | lav lvs | ✓ | ||
mi | Maori | mri mao | |||
mk | Macedonian | mkd mac | ✓ | ✓ | |
ml | Malayalam | mal | |||
mn | Mongolian | mon khk | ✓ | ✓ | |
mr | Marathi | mar | ✓ | ✓ | |
ms | Malay | msa zsm | ✓ | ✓ | ✓ |
my | Burmese (Myanmar) | mya bur | ✓ | ✓ | |
nb | Norwegian Bokmål (Norwegian) | nob nor | ✓ | ||
nl | Dutch | nld dut | ✓ | ✓ | ✓ |
nn | Norwegian Nynorsk (Nynorsk) | nno | |||
or | Odia (Oriya) | ori ory | |||
pa | Punjabi (Panjabi) | pan | |||
pl | Polish | pol | ✓ | ✓ | ✓ |
pt | Portuguese | por | ✓ | ✓ | ✓ |
ro | Romanian | ron rum | ✓ | ✓ | |
ru | Russian | rus | ✓ | ✓ | ✓ |
si | Sinhala (Sinhalese) | sin | |||
sk | Slovak | slk slo | ✓ | ||
sl | Slovene (Slovenian) | slv | ✓ | ||
sn | Shona | sna | |||
so | Somali | som | |||
sq | Albanian | sqi als | ✓ | ||
sr | Serbian | srp | ✓ | ||
st | Southern Sotho (Sotho) | sot | |||
sv | Swedish | swe | ✓ | ✓ | |
sw | Swahili | swa swh | ✓ | ||
ta | Tamil | tam | ✓ | ✓ | |
te | Telugu | tel | ✓ | ||
th | Thai | tha | ✓ | ✓ | ✓ |
tl | Tagalog (Filipino) | tgl fil | ✓ | ✓ | ✓ |
tn | Tswana | tsn | |||
tr | Turkish | tur | ✓ | ✓ | ✓ |
ts | Tsonga | tso | |||
uk | Ukrainian | ukr | ✓ | ✓ | |
ur | Urdu | urd | ✓ | ||
vi | Vietnamese | vie | ✓ | ✓ | ✓ |
xh | Xhosa | xho | |||
yo | Yoruba | yor | |||
yue | Cantonese (Yue Chinese) | ✓ | ✓ | ||
zh | Chinese (Mandarin) | zho cmn | ✓ | ✓ | ✓ |
zu | Zulu | zul |
Alignment: send the language when you know it
When language is omitted, alignment infers it from the reference text alone. Languages that share a writing system can be confused this way. Cantonese written in Chinese characters, for example, may be read as Mandarin, and only a few words line up with the audio. Sending language=yue explicitly avoids this kind of error from automatic language inference.
If you know the language, send it. Use und only for mixed-language or romanized text, or for a language that has no code.
Examples
Transcribe Cantonese audio:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/transcriptions" \
-F file=@cantonese.wav -F language=yue -F response_format=jsonAlign a known Cantonese transcript with the same audio:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
-F file=@cantonese.wav -F text="今日天氣好好。" -F language=yueNormalize written text into spoken Traditional Chinese:
curl --fail-with-body "$EDGESPEAK_BASE_URL/text/normalizations" \
-H "Content-Type: application/json" \
-d '{"text":"8:15","language":"zh-TW","mode":"tn","top_k":1}'Example output (excerpt):
{
"task": "normalize",
"language": "zh",
"mode": "tn",
"alternatives": [{ "rank": 0, "text": "八點十五分" }]
}The response reports the resolved code zh; the zh-TW tag only selected Traditional characters.
An unsupported transcription language returns a stable error code. Branch on error.code, not on the message:
{"error":{"message":"unsupported language: 'tlh'","type":"invalid_request_error","code":"asr_language_unsupported","param":"language"}}message echoes the value you sent for diagnostics only.