Extract speaker embeddings
POST /v1/speaker/embeddings
Extracts one 256-dimensional embedding per uploaded sample. Repeat the file part for several samples. Each sample must be at most 30 seconds and contain at least 4 seconds of effective single-speaker speech.
Guide and examplesTranscription and speakersAll endpoints
Request body
multipart/form-dataRequired
filefile[]RequiredRepeat this part once per sample.modelstringSpeaker embedding model. Defaults to EdgeSpeak/Lattice-1. Embeddings are only comparable when they come from the same model.Responses
200One entry per uploaded sample, in upload order. Full 256-value vectors are elided in this example.objectstringmodelstringdataobject[]indexintegerembeddingnumber[]256-dimensional vector.effective_speech_secondsnumber{
"object": "list",
"model": "EdgeSpeak/Lattice-1",
"data": [
{
"index": 0,
"embedding": [
0.0123,
-0.0456
],
"effective_speech_seconds": 5.2
},
{
"index": 1,
"embedding": [
0.0311,
-0.0092
],
"effective_speech_seconds": 6.1
}
]
}400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.401Invalid or missing API key.403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.413Request body exceeds 512 MiB.503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.