Transcribe an audio file

POST /v1/audio/transcriptions

Transcribes an uploaded audio or video file. json returns text with duration, verbose_json returns a structured transcript, text returns plain text, and diarized_json also runs speaker diarization. Word or segment granularities require verbose or diarized output. stream=true is desktop-gateway only and returns SSE text deltas with JSON output and no granularities; the headless listener refuses it with 400.

Guide and examplesTranscription and speakersLocal Gateway APIAll endpoints

Request body

multipart/form-dataRequired

filefileRequiredAudio or video file. The general request-body limit is 512 MiB.
modelstringOmit to use the service's configured default.
languagestringOptional; omit it to detect the language from the audio. Validated against the supported transcription languages. A code outside that set returns 400 asr_language_unsupported with param=language. See /docs/languages#transcription for supported languages.
response_formatstringOutput mode.json | verbose_json | text | diarized_json
streambooleanDesktop gateway only, with JSON output and no timestamp granularities.
timestamp_granularities[]string[]Requires verbose or diarized output.word | segment
word_timestamp_enabledbooleanControls word timestamps directly.
candidatesintegerN-best output. 1 is normal; 2-8 require non-streaming verbose_json without word timestamps. A conflicting repeated value returns 400.1–8
semantic_sentence_enabledbooleanControls semantic processing. An explicit false conflicts with length or margin fields and returns 400.
min_charsintegerDefault length 12. Length constraints are off by default; supplying a length field enables them and requests semantic processing.
max_charsintegerDefault length 42.
start_marginnumberSeconds; default margin 0.2.
end_marginnumberSeconds; default margin 0.2.

Responses

200The transcript, in the shape selected by response_format. A streaming request returns SSE transcript.text.delta events followed by transcript.text.done; an HTTP 200 does not by itself prove that a stream completed.

TranscriptJson

textstring
startnumber
durationnumber
languagestring
usageobjectAudio duration accounting for transcription and alignment.
typestringduration
secondsnumberSeconds of audio processed. For a windowed alignment this is the processed window, not the full media duration.

TranscriptVerbose

taskstringtranscribe
durationnumber
languagestring
textstring
segmentsobject[]
idintegerSequential index of the segment within the result.
startnumberSegment start in seconds.
endnumberSegment end in seconds.
textstringSegment text.
speakerstring | nullSpeaker label, or null when no speaker was assigned.
wordsobject[]Word timings. Present only when word timestamps were requested.
word_timestamps_unavailableobject[]Segments whose alignment failed. Their text is kept but they carry no words, and their times are the original segment window. Each item is {start, end} in seconds. Omitted when every segment aligned.
startnumber
endnumber
usageobjectAudio duration accounting for transcription and alignment.
typestringduration
secondsnumberSeconds of audio processed. For a windowed alignment this is the processed window, not the full media duration.

TranscriptDiarized

textstring
segmentsobject[]
typestringtranscript.text.segment
idstring
startnumber
endnumber
textstring
speakerstring | null
wordsobject[]
word_timestamps_unavailableobject[]Segments whose alignment failed. Their text is kept but they carry no words, and their times are the original segment window. Each item is {start, end} in seconds. Omitted when every segment aligned.
startnumber
endnumber
usageobjectAudio duration accounting for transcription and alignment.
typestringduration
secondsnumberSeconds of audio processed. For a windowed alignment this is the processed window, not the full media duration.
JSON
{
  "text": "Hello world.",
  "start": 0,
  "duration": 2.5,
  "language": "en",
  "usage": {
    "type": "duration",
    "seconds": 2.5
  }
}
400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.
401Invalid or missing API key.
403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.
413Request body exceeds 512 MiB.
503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.