Languages

Use one set of short language codes across the APIs.

Use the same short, lowercase code such as zh, yue, en or ja for transcription, alignment, text normalization, speech, Realtime, MCP tools and CLI options. Which other spellings each API accepts, and how it normalizes them, is described below; the speech example filter and the voice library are exceptions.

This page covers the language of your audio or text. It is unrelated to the interface language of the app or the CLI, and to language models.

Code format

These rules apply to transcription, alignment, text normalization, speech synthesis and Realtime, and to the MCP tools and CLI options that call them. See Exceptions for the two places that do not normalize.

  • Recommended form: the two-letter ISO 639-1 code. The only three-letter code is Cantonese, yue.
  • Case and separators: matching ignores case and surrounding whitespace, and _ works like -. ZH, en_US and ja are all valid.
  • Region and script subtags: full BCP-47 tags are accepted and resolve to their language: en-US and en-GB → en, pt-BR → pt, zh-CN, zh-Hans, zh-TW and zh-Hant-TW → zh.
  • Other spellings: ISO 639-3 codes (eng, deu, cmn, fil) and English names (English, Mandarin, Cantonese) are also recognized. Prefer the short code in new integrations.
  • Responses: the language field returned by transcription, alignment and text normalization is the resolved code, for example zh or yue.
You sendResolves to
zh, zh-CN, zh-Hans, zh-TW, zh-HK, zh-Hant, cmn, Mandarinzh
yue, zh-yue, Cantoneseyue
en, en-US, en-GB, eng, Englishen
pt, pt-BR, pt-PTpt
tl, fil, Filipinotl
nb, nob, nor, Norwegiannb
zh-HK and zh-TW mean Chinese (Mandarin), not Cantonese. For Cantonese audio or text, send yue.

Chinese script is not a separate language. Transcription and alignment treat zh-Hant as zh. In text normalization, an explicit Chinese tag decides the script of the readings it generates in tn mode: zh-Hant, or a TW, HK or MO region, gives Traditional characters. A script subtag wins over the region, so zh-Hans-TW stays Simplified. Text kept from the input is never converted, and itn output is not affected. Translation keeps Traditional Chinese as its own target, zh-hant.

Omitting it, und and auto

The code format is shared. What happens when you leave language out, and which special values exist, depends on the capability:

CapabilityOmittedundauto
TranscriptionDetects the language from the audioRejectedRejected
AlignmentInfers the language from the reference text; if it cannot, aligns without a languageNo language information: each word is matched against every language, and normalization reads segment by segmentRejected
Text normalizationRejected, language is requiredNo language information: readings are chosen segment by segment from the textRejected
SpeechChosen automaticallyNot a speech valueSame as omitting it

und is the BCP-47 tag for "undetermined". UND and und-Latn are treated as und.

Parameters by API

APIParameterRequiredWhen omittedUnsupported value
POST /v1/audio/transcriptionslanguage (multipart)NoDetected from audio400 asr_language_unsupported, param=language
POST /v1/audio/alignmentslanguage (multipart)NoInferred from the reference text400 bad_request for a malformed tag; 422 language_unsupported when the selected model does not cover the language; both with param=language
POST /v1/text/normalizationslanguage (JSON)Yes400 normalization_invalid_request400 normalization_invalid_request for an unrecognized tag; 422 normalization_language_unsupported for a language without normalization rules
POST /v1/audio/speechlanguage, segments[].languageNoChosen automatically422 language_unsupported
POST /v1/audio/voiceslanguage (multipart)NoSaved as zh-CN—
GET /v1/audio/speech/exampleslanguage (query)NoNo filter—
WS /v1/realtimesession.audio.input.transcription.languageNoDetected from audio—
WS /v1/realtimesession.audio.output.languageNoUses the detected spoken language—

In Realtime, session.audio.output.language is the conversation language: built-in voices use it to choose a reference profile until a spoken language has been detected, after which the detected language takes precedence. /v1/text/segmentations and the speaker endpoints do not take a language. Voice library entries store the tag as sent; see Exceptions.

MCP tools and CLI options pass the same values to the same capabilities:

MCP toolCLI optionCapability
edgespeak_transcribe, edgespeak_transcribe_file: languageedgespeak-cli transcribe --languageTranscription
edgespeak_align: languageedgespeak-cli align --languageAlignment
edgespeak_create_speech: languageedgespeak-cli speech --languageSpeech
edgespeak_add_voice: languageedgespeak-cli voices add --languageVoice library, default zh-CN
edgespeak_speech_examples: languageedgespeak-cli speech-examples --languageExample filter

edgespeak-cli transcribe --language needs a reachable gateway, either the desktop app or edgespeak-cli serve. Without one, the command exits with an error instead of ignoring the option.

Exceptions

  • Speech example filter (GET /v1/audio/speech/examples, edgespeak_speech_examples, edgespeak-cli speech-examples): only checks whether the first subtag is zh. Every other value, including cmn, Mandarin, yue and zh with surrounding spaces, returns English examples. Send zh or en.
  • Voice library (POST /v1/audio/voices, edgespeak_add_voice, edgespeak-cli voices add): trims surrounding whitespace and saves the tag as sent. It becomes the voice's locale and the key of names; English is not converted to en or en-US. Send a locale such as zh-CN or en-US.

Common languages

With the default alignment model, alignment covers every language in this table.

LanguageCodeAlso acceptedTranscriptionNormalizationTranslation
Chinese (Mandarin)zhzh-CN, zh-TW, zh-Hans, zh-Hant, cmn✓✓✓
Cantoneseyuezh-yue✓—✓
Englishenen-US, en-GB✓✓✓
Japanesejaja-JP✓✓✓
Koreankoko-KR✓✓✓
Frenchfrfr-FR, fr-CA✓✓✓
Germandede-DE✓✓✓
Spanisheses-ES, es-MX✓✓✓
Portugueseptpt-BR, pt-PT✓✓✓
Russianruru-RU✓✓✓
Arabicarar-SA✓✓✓
Italianitit-IT✓✓✓
Vietnamesevivi-VN✓✓✓
Thaithth-TH✓✓✓
Indonesianidid-ID✓✓✓
Hindihihi-IN✓✓✓

Speech support depends on the model; see Speech.

Supported languages by capability

The capabilities cover different sets. Check the one you call rather than assuming every language works everywhere.

Transcription

30 languages: zh en yue ar de fr es pt id it ko ru th vi ja tr hi ms nl sv da fi pl cs tl fa el hu mk ro.

The same list applies to every transcription model.

Alignment

ModelLanguages
EdgeSpeak/Lattice-2 (default)Every language in the catalog, mixed-language text, und, and other well-formed tags such as nan or wuu, which have no dictionary pronunciations and perform worse than und
EdgeSpeak/Lattice-1zh, en, de only. und and tags outside the catalog return 400 bad_request; other catalog languages return 422 language_unsupported

Tags outside the catalog keep their script and region subtags, for example nan-Hant. For how alignment applies text normalization, see Forced alignment.

Text normalization

55 languages, the same for tn and itn: ar az bg bs ca cs da de el en es et fa fi fr he hi hr hu hy id is it ja ka kk km ko lo lt lv mk mn mr ms my nb nl pl pt ro ru sk sl sq sr sv sw ta th tl tr uk vi zh.

Cantonese yue has no normalization rules. The running list and the categories per language are in the EdgeSpeak/Skylark entry of GET /v1/models: x_edgespeak.normalization.languages_by_mode and classes_by_language.

Speech

Each speech model supports its own set. Read supported_languages for the model in GET /v1/models; when x_edgespeak.speech_language.open_set is true, the model also accepts ISO 639 codes outside that list. Region and script variants of a listed language, such as zh-TW or en-US, are accepted, and auto always is.

ModelLanguages
k2-fsa/OmniVoiceOpen set: every catalog language plus other ISO 639 codes
Qwen/Qwen3-TTS-*zh en de fr es pt it ko ru ja
IndexTeam/IndexTTS-2.5zh en ja es ar
FireRedTeam/FireRedTTS3zh yue en de fr es it pt ru ja ko ar cs nl fi el hi id pl ro th tr uk vi, plus 21 Chinese dialects written zh-x-<name>, such as zh-x-sichuan
FireRedTeam/FireRedTTS3-InstructSame 24 languages, without the dialects
openbmb/VoxCPM2zh en de ja
BreezeBlue/Breeze-TTS-2zh en

FireRedTTS3-Instruct, VoxCPM2 and Breeze-TTS-2 do not take a language input; for them language selects the matching reference audio and is still checked against the list. List every model's languages from a running service:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/models" \
  | jq '.data[] | select(.supported_languages) | {id, supported_languages}'

Translation

38 target languages: the 36 marked in the catalog, plus Traditional Chinese zh-hant and Uyghur ug. Translation tags are stored in lowercase, as in the target_langs field of JSON exports. See Translation requirements for how translation works.

Language catalog

EdgeSpeak recognizes 82 languages. Any code, ISO 639-3 code or English name in this table is accepted as input; the default alignment model covers all of them. The last three columns show which languages the other capabilities support.

CodeLanguageISO 639-3TranscriptionNormalizationTranslation
afAfrikaansafr
arArabicara arb✓✓✓
azAzerbaijaniaze azj✓
beBelarusianbel
bgBulgarianbul✓
bnBengali (Bangla)ben✓
boTibetanbod tib✓
bsBosnianbos✓
caCatalancat✓
csCzechces cze✓✓✓
cyWelshcym wel
daDanishdan✓✓
deGermandeu ger✓✓✓
elGreekell gre✓✓
enEnglisheng✓✓✓
esSpanish (Castilian)spa✓✓✓
etEstonianest ekk✓
euBasqueeus
faPersian (Farsi)fas pes✓✓✓
fiFinnishfin✓✓
frFrenchfra fre✓✓✓
gaIrishgle
guGujaratiguj✓
heHebrewheb✓✓
hiHindihin✓✓✓
hrCroatianhrv✓
huHungarianhun✓✓
hyArmenianhye✓
idIndonesianind✓✓✓
isIcelandicisl ice✓
itItalianita✓✓✓
jaJapanesejpn✓✓✓
kaGeorgiankat geo✓
kkKazakhkaz✓✓
kmKhmer (Cambodian)khm✓✓
knKannadakan
koKoreankor✓✓✓
lgGanda (Luganda)lug
loLaolao✓
ltLithuanianlit✓
lvLatvianlav lvs✓
miMaorimri mao
mkMacedonianmkd mac✓✓
mlMalayalammal
mnMongolianmon khk✓✓
mrMarathimar✓✓
msMalaymsa zsm✓✓✓
myBurmese (Myanmar)mya bur✓✓
nbNorwegian Bokmål (Norwegian)nob nor✓
nlDutchnld dut✓✓✓
nnNorwegian Nynorsk (Nynorsk)nno
orOdia (Oriya)ori ory
paPunjabi (Panjabi)pan
plPolishpol✓✓✓
ptPortuguesepor✓✓✓
roRomanianron rum✓✓
ruRussianrus✓✓✓
siSinhala (Sinhalese)sin
skSlovakslk slo✓
slSlovene (Slovenian)slv✓
snShonasna
soSomalisom
sqAlbaniansqi als✓
srSerbiansrp✓
stSouthern Sotho (Sotho)sot
svSwedishswe✓✓
swSwahiliswa swh✓
taTamiltam✓✓
teTelugutel✓
thThaitha✓✓✓
tlTagalog (Filipino)tgl fil✓✓✓
tnTswanatsn
trTurkishtur✓✓✓
tsTsongatso
ukUkrainianukr✓✓
urUrduurd✓
viVietnamesevie✓✓✓
xhXhosaxho
yoYorubayor
yueCantonese (Yue Chinese)✓✓
zhChinese (Mandarin)zho cmn✓✓✓
zuZuluzul

Alignment: send the language when you know it

When language is omitted, alignment infers it from the reference text alone. Languages that share a writing system can be confused this way. Cantonese written in Chinese characters, for example, may be read as Mandarin, and only a few words line up with the audio. Sending language=yue explicitly avoids this kind of error from automatic language inference.

If you know the language, send it. Use und only for mixed-language or romanized text, or for a language that has no code.

Examples

Transcribe Cantonese audio:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/transcriptions" \
  -F file=@cantonese.wav -F language=yue -F response_format=json

Align a known Cantonese transcript with the same audio:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F file=@cantonese.wav -F text="今日天氣好好。" -F language=yue

Normalize written text into spoken Traditional Chinese:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/text/normalizations" \
  -H "Content-Type: application/json" \
  -d '{"text":"8:15","language":"zh-TW","mode":"tn","top_k":1}'

Example output (excerpt):

JSON
{
  "task": "normalize",
  "language": "zh",
  "mode": "tn",
  "alternatives": [{ "rank": 0, "text": "八點十五分" }]
}

The response reports the resolved code zh; the zh-TW tag only selected Traditional characters.

An unsupported transcription language returns a stable error code. Branch on error.code, not on the message:

JSON
{"error":{"message":"unsupported language: 'tlh'","type":"invalid_request_error","code":"asr_language_unsupported","param":"language"}}

message echoes the value you sent for diagnostics only.