Transcribe an audio file
POST /v1/audio/transcriptions
Transcribes an uploaded audio or video file. json returns text with duration, verbose_json returns a structured transcript, text returns plain text, and diarized_json also runs speaker diarization. Word or segment granularities require verbose or diarized output. stream=true is desktop-gateway only and returns SSE text deltas with JSON output and no granularities; the headless listener refuses it with 400.
Guide and examplesTranscription and speakersLocal Gateway APIAll endpoints
Request body
multipart/form-dataRequired
filefileRequiredAudio or video file. The general request-body limit is 512 MiB.modelstringOmit to use the service's configured default.languagestringOptional; omit it to detect the language from the audio. Validated against the supported transcription languages. A code outside that set returns 400 asr_language_unsupported with param=language. See /docs/languages#transcription for supported languages.response_formatstringOutput mode.streambooleanDesktop gateway only, with JSON output and no timestamp granularities.timestamp_granularities[]string[]Requires verbose or diarized output.word_timestamp_enabledbooleanControls word timestamps directly.candidatesintegerN-best output. 1 is normal; 2-8 require non-streaming verbose_json without word timestamps. A conflicting repeated value returns 400.semantic_sentence_enabledbooleanControls semantic processing. An explicit false conflicts with length or margin fields and returns 400.min_charsintegerDefault length 12. Length constraints are off by default; supplying a length field enables them and requests semantic processing.max_charsintegerDefault length 42.start_marginnumberSeconds; default margin 0.2.end_marginnumberSeconds; default margin 0.2.Responses
200The transcript, in the shape selected by response_format. A streaming request returns SSE transcript.text.delta events followed by transcript.text.done; an HTTP 200 does not by itself prove that a stream completed.TranscriptJson
textstringstartnumberdurationnumberlanguagestringusageobjectAudio duration accounting for transcription and alignment.typestringsecondsnumberSeconds of audio processed. For a windowed alignment this is the processed window, not the full media duration.TranscriptVerbose
taskstringdurationnumberlanguagestringtextstringsegmentsobject[]idintegerSequential index of the segment within the result.startnumberSegment start in seconds.endnumberSegment end in seconds.textstringSegment text.speakerstring | nullSpeaker label, or null when no speaker was assigned.wordsobject[]Word timings. Present only when word timestamps were requested.word_timestamps_unavailableobject[]Segments whose alignment failed. Their text is kept but they carry no words, and their times are the original segment window. Each item is {start, end} in seconds. Omitted when every segment aligned.startnumberendnumberusageobjectAudio duration accounting for transcription and alignment.typestringsecondsnumberSeconds of audio processed. For a windowed alignment this is the processed window, not the full media duration.TranscriptDiarized
textstringsegmentsobject[]typestringidstringstartnumberendnumbertextstringspeakerstring | nullwordsobject[]word_timestamps_unavailableobject[]Segments whose alignment failed. Their text is kept but they carry no words, and their times are the original segment window. Each item is {start, end} in seconds. Omitted when every segment aligned.startnumberendnumberusageobjectAudio duration accounting for transcription and alignment.typestringsecondsnumberSeconds of audio processed. For a windowed alignment this is the processed window, not the full media duration.{
"text": "Hello world.",
"start": 0,
"duration": 2.5,
"language": "en",
"usage": {
"type": "duration",
"seconds": 2.5
}
}400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.401Invalid or missing API key.403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.413Request body exceeds 512 MiB.503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.