Get a speaker activity timeline
POST /v1/speaker/diarizations
Returns who spoke when, without transcription. Speaker endpoints accept uploaded bytes rather than arbitrary local paths, and num_speakers is clustering guidance, not an identity claim.
Guide and examplesTranscription and speakersAll endpoints
Request body
multipart/form-dataRequired
filefileRequiredAudio or video file. The general request-body limit is 512 MiB.num_speakersintegerOptional clustering guidance.Responses
200The speaker timeline. No text is returned.durationnumberspeakersstring[]segmentsobject[]startnumberendnumberspeakerstring{
"duration": 10,
"speakers": [
"SPEAKER_00",
"SPEAKER_01"
],
"segments": [
{
"start": 0.4,
"end": 3.2,
"speaker": "SPEAKER_00"
},
{
"start": 4,
"end": 8.1,
"speaker": "SPEAKER_01"
}
]
}400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.401Invalid or missing API key.403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.413Request body exceeds 512 MiB.503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.