Extract speaker embeddings

POST /v1/speaker/embeddings

Extracts one 256-dimensional embedding per uploaded sample. Repeat the file part for several samples. Each sample must be at most 30 seconds and contain at least 4 seconds of effective single-speaker speech.

Guide and examplesTranscription and speakersAll endpoints

Request body

multipart/form-dataRequired

filefile[]RequiredRepeat this part once per sample.
modelstringSpeaker embedding model. Defaults to EdgeSpeak/Lattice-1. Embeddings are only comparable when they come from the same model.Default "EdgeSpeak/Lattice-1"

Responses

200One entry per uploaded sample, in upload order. Full 256-value vectors are elided in this example.
objectstringlist
modelstring
dataobject[]
indexinteger
embeddingnumber[]256-dimensional vector.
effective_speech_secondsnumber
JSON
{
  "object": "list",
  "model": "EdgeSpeak/Lattice-1",
  "data": [
    {
      "index": 0,
      "embedding": [
        0.0123,
        -0.0456
      ],
      "effective_speech_seconds": 5.2
    },
    {
      "index": 1,
      "embedding": [
        0.0311,
        -0.0092
      ],
      "effective_speech_seconds": 6.1
    }
  ]
}
400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.
401Invalid or missing API key.
403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.
413Request body exceeds 512 MiB.
503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.