Transcription and speakers
Transcribe an audio file
POST /v1/audio/transcriptions
Replace hello.wav with your own audio or video file. The multipart file field is required; omitting model uses the service's configured default. Use response_format=json when you only need text; see the next example if you need timestamps or speaker labels.
Request:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/transcriptions" \
-F model=EdgeSpeak/Lattice-2 \
-F file=@hello.wav -F response_format=jsonExample output:
{
"text": "Hello world.",
"start": 0.0,
"duration": 2.5,
"language": "en",
"usage": {"type": "duration", "seconds": 2.5}
}Setting response_format to text returns plain text, verbose_json returns a structured transcript, and diarized_json includes speaker diarization. See transcript field definitions and optional parameters for the full request parameters.
The language parameter is optional; omit it to detect the language from the audio. It is validated against the supported transcription languages, and codes follow the shared language code format. Passing an unsupported language code returns HTTP 400 asr_language_unsupported, with param=language identifying the invalid field. The message in the response echoes the value you sent for diagnostics; clients should branch on error.code rather than parsing error text.
Get text, timing and speakers together
Specify response_format=diarized_json when you need text, timestamps and “who said what” together. The returned speaker labels represent voice clusters within this recording, not real personal identities. Timings and labels below are illustrative.
Request:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/transcriptions" \
-F model=EdgeSpeak/Lattice-2 \
-F file=@hello.wav -F response_format=diarized_json \
-F "timestamp_granularities[]=word"Example output:
{
"text": "Hello world.",
"segments": [
{
"type": "transcript.text.segment",
"id": "seg_0",
"start": 0.2,
"end": 1.4,
"text": "Hello world.",
"speaker": "SPEAKER_00",
"words": [
{
"word": "Hello",
"start": 0.2,
"end": 0.6
},
{
"word": "world.",
"start": 0.7,
"end": 1.4
}
]
}
],
"usage": {
"type": "duration",
"seconds": 2.5
}
}The CLI transcribe --diarize command calls this endpoint under the hood; see API reuse for parameter mapping and differences from standalone defaults. The candidates parameter (values 2–8) applies only to non-streaming verbose_json and cannot be used alongside word timestamps.
Read transcription events
This mode is supported only by the desktop gateway (the Headless service currently returns HTTP 400 for this request). When calling it, set stream=true and response_format=json, and do not specify timestamp granularities. Use curl -N to disable client buffering.
Request:
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/transcriptions" \
-F model=EdgeSpeak/Lattice-2 \
-N -F file=@hello.wav -F response_format=json -F stream=trueExample output:
event: transcript.text.delta
data: {"type":"transcript.text.delta","delta":"Hello world."}
event: transcript.text.done
data: {"type":"transcript.text.done","text":"Hello world.","usage":{"type":"duration","seconds":2.5}}Clients should append delta text increments to the display and use the complete text in the done event as the final result. Also handle error events and unexpected connection disconnects.
Get a speaker activity timeline
POST /v1/speaker/diarizations
Call this endpoint when you only need timeline intervals of “who spoke when” without transcription text. Upload an audio file; the optional num_speakers parameter (1–32) guides the number of speaker clusters.
Request:
curl --fail-with-body "$EDGESPEAK_BASE_URL/speaker/diarizations" \
-F file=@meeting.wav -F num_speakers=2Example output:
{
"duration": 10.0,
"speakers": [
"SPEAKER_00",
"SPEAKER_01"
],
"segments": [
{
"start": 0.4,
"end": 3.2,
"speaker": "SPEAKER_00"
},
{
"start": 4.0,
"end": 8.1,
"speaker": "SPEAKER_01"
}
]
}The response contains no transcription text. The num_speakers parameter serves only as clustering guidance, and the returned speaker labels do not correspond to real personal identities.
Extract speaker embeddings
POST /v1/speaker/embeddings
Batch upload audio samples using multiple file fields. Each sample must be at most 30 seconds and contain at least 4 seconds of effective single-speaker speech. Save the complete response for subsequent similarity comparisons; the example uses jq to extract and print only vector lengths instead of printing all 256 numbers. Running this example requires jq.
Request:
curl --fail-with-body "$EDGESPEAK_BASE_URL/speaker/embeddings" \
-F model=EdgeSpeak/Lattice-1 \
-F file=@speaker-a.wav -F file=@speaker-b.wav \
-o embeddings.json
jq '.data | map({index, dimensions: (.embedding | length), effective_speech_seconds})' embeddings.jsonExample output:
[
{
"index": 0,
"dimensions": 256,
"effective_speech_seconds": 5.2
},
{
"index": 1,
"dimensions": 256,
"effective_speech_seconds": 6.1
}
]The output above is processed by jq; the raw HTTP response structure is {object:"list",model,data:[{index,embedding,effective_speech_seconds}]}. The optional model parameter defaults to EdgeSpeak/Lattice-1. Retain the saved embeddings.json for the next step.
Compare the two embeddings
POST /v1/speaker/similarity
Continuing from the previous step, construct the request body using the actual returned vectors. Both vectors being compared must contain 256-dimensional finite numerical values with a non-zero norm.
Request:
jq '{model, embedding_a: .data[0].embedding, embedding_b: .data[1].embedding}' \
embeddings.json > similarity-request.json
curl --fail-with-body "$EDGESPEAK_BASE_URL/speaker/similarity" \
-H "Content-Type: application/json" -d @similarity-request.jsonExample output:
{
"object": "speaker_similarity",
"model": "EdgeSpeak/Lattice-1",
"similarity": 0.82
}The similarity score shown is illustrative; actual cosine similarity ranges between -1 and 1. The endpoint accepts neither threshold nor match parameters; your application decides how to use the score.
Continue: Forced alignment · API index