Align known text with audio

POST /v1/audio/alignments

Locates known words in audio rather than recognizing unknown speech. Send file and matching text; the desktop gateway also accepts an absolute local text_path, while Headless requires text. Timestamps stay relative to the original audio: duration is the full media duration and usage.seconds measures the processed window. Sending the progress: 1 header returns an NDJSON progress stream whose result line carries task: "align"; parse each line independently, not as SSE. See /docs/api-align#align and /docs/api#alignment-options. A failed alignment returns 422 with an error.alignment object instead of a result; search_effort chooses the search range and can retry once with a wider one.

Guide and examplesForced alignmentAll endpoints

Parameters

progressstringSet to 1 for an NDJSON alignment progress stream. Ordinary requests return one JSON response.1

Request body

multipart/form-dataRequired

filefileRequiredAudio file.
textstringTranscript to align. Required on Headless.
text_pathstringDesktop gateway only: absolute local path to the transcript.
modelstringOptional alignment model. The default EdgeSpeak/Lattice-2 covers every language in the catalog, mixed-language text and und. EdgeSpeak/Lattice-1 covers Chinese, English and German only and must be requested explicitly. A transcription or speech model returns 400 model_capability_unsupported, and an unknown ID returns 404 model_not_found; both carry param: "model".
languagestringOptional hint. Omit it to let the runtime detect the language, send a code such as zh or yue to pin one, or send und to skip detection and assume no language. EdgeSpeak/Lattice-1 rejects und and out-of-list codes with 400 bad_request, and a language it does not cover with 422 language_unsupported. See /docs/languages#alignment for supported languages.
startnumberWindow start in seconds. Supply together with end, with 0 <= start < end and within the audio.
endnumberWindow end in seconds.
search_effortstringSearch range. standard searches once with the standard range; an empty value ("") means the same. extended searches once with a wider range, slower and using more memory. auto runs standard and, only if it fails with error.alignment.retry_recommended: true, runs extended once. Other values return 400 bad_request with param: "search_effort".standard | extended | auto | Default "standard"
protected_terms[]string[]Terms that must align as a single word rather than being split, such as product names or jargon. Repeat the field, or send a JSON array in protected_terms.
text_normalizationstringA JSON string inside multipart. Omitting the whole field enables normalization; {"enabled":false} or an empty object disables it. Fields are enabled, top_k (default 8, range 1-16), classes, apply_dictionary (default false), dictionary_ids and pins with {source_start_byte, source_end_byte, rank} over the original UTF-8 transcript, where rank: null preserves the source. When disabled, do not supply top_k, classes or non-empty pins.
audio_trackintegerZero-based decodable audio track. A conflicting repeated value returns 400.
audio_channelstringmix or a zero-based channel index. A conflicting repeated value returns 400.
semantic_sentence_enabledbooleanControls semantic processing. An explicit false conflicts with length or margin fields and returns 400.
min_charsintegerDefault length 12. Length constraints are off by default; supplying a length field enables them and requests semantic processing.
max_charsintegerDefault length 42.
start_marginnumberSeconds; default margin 0.2.
end_marginnumberSeconds; default margin 0.2.

Responses

200The alignment result. With progress: 1 the body is NDJSON: progress lines, a {"stage":"extended_search"} line when search_effort=auto starts the wider search (progress then restarts from 0), and a last line that is either the result (task: "align") or the same {"error": {...}} envelope as the 422 body. The HTTP status of the stream stays 200; progress 1.0 alone does not prove success, and a stream that closes without a result or error line is a failure.
taskstringalign
durationnumberFull media duration.
languagestringCanonical language code, when the runtime resolved one.
language_candidatesstring[]Ordered candidate languages the alignment actually used, preferred first. Comes from the request language or from the reference text; it routes pronunciation and is not an acoustic check of the audio language. Omitted when the candidate list is empty, for example when the request passes und or no language can be resolved.
textstring
segmentsobject[]
idintegerSequential index of the segment within the result.
startnumberSegment start in seconds.
endnumberSegment end in seconds.
textstringSegment text.
speakerstring | nullSpeaker label, or null when no speaker was assigned.
wordsobject[]Word timings. Present only when word timestamps were requested.
normalization_regionsobject[]Regions and candidates produced by text normalization. Word-level provenance carries source_start_byte, source_end_byte and normalizations[], relative to the original UTF-8 transcript.
search_effort_usedstringThe search that actually ran. Present when the service reports it.standard | extended
search_retriedbooleantrue when search_effort=auto actually ran the wider search after the first attempt failed. Omitted otherwise.
usageobjectAudio duration accounting for transcription and alignment.
typestringduration
secondsnumberSeconds of audio processed. For a windowed alignment this is the processed window, not the full media duration.
JSON
{
  "task": "align",
  "duration": 2.5,
  "text": "Hello world.",
  "segments": [
    {
      "id": 0,
      "start": 0.2,
      "end": 1.4,
      "text": "Hello world.",
      "words": [
        {
          "word": "Hello",
          "start": 0.2,
          "end": 0.6,
          "score": 0.98
        },
        {
          "word": "world.",
          "start": 0.7,
          "end": 1.4,
          "score": 0.96
        }
      ]
    }
  ],
  "usage": {
    "type": "duration",
    "seconds": 2.5
  }
}
400Invalid request. Fix the request before retrying; use error.code for program logic and error.param to locate the input.
401Invalid or missing API key.
403Host or origin is not allowed, or the license was rejected. Inspect error.code to tell them apart.
404Unknown model ID (model_not_found).
413Request body exceeds 512 MiB.
422alignment_failed: alignment ran but could not be completed. alignment_search_budget_exceeded: a wider search would need more memory than this computer can spare, so it did not run. Both carry error.alignment with reason (no_path, no_words, collapsed, search_incomplete), search_effort_used, search_path (whole_audio, streaming), search_limit (none, budget, graph_size, estimate_overflow, streaming), retry_recommended (always present) and estimated_extra_bytes (an estimate of the extra memory a wider search needs, not a cap); unknown fields are omitted. Failed alignments are not counted in usage. The same status also covers language_unsupported (change the model, not the language code) and normalization_class_unsupported.
JSON
{
  "error": {
    "message": "Alignment could not be completed; the search stopped before the end of the text.",
    "type": "invalid_request_error",
    "code": "alignment_failed",
    "alignment": {
      "reason": "search_incomplete",
      "search_effort_used": "standard",
      "search_path": "whole_audio",
      "search_limit": "none",
      "retry_recommended": true,
      "estimated_extra_bytes": 3221225472
    }
  }
}
503A required local model is still being prepared (model_downloading) or the service is busy (service_busy). Honor Retry-After when present.