Forced alignment

Align known text with audio

POST /v1/audio/alignments

Pass file and matching text to locate timestamps for known text in audio rather than recognizing unknown speech. The example disables text normalization to preserve the original spelling; omitting this configuration enables text normalization by default.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@hello.wav -F text="Hello world." -F language=en \
  -F 'text_normalization={"enabled":false}'

Example output:

JSON
{
  "task": "align",
  "duration": 2.5,
  "text": "Hello world.",
  "segments": [
    {
      "id": 0,
      "start": 0.2,
      "end": 1.4,
      "text": "Hello world.",
      "words": [
        {
          "word": "Hello",
          "start": 0.2,
          "end": 0.6,
          "score": 0.98
        },
        {
          "word": "world.",
          "start": 0.7,
          "end": 1.4,
          "score": 0.96
        }
      ]
    }
  ],
  "usage": {
    "type": "duration",
    "seconds": 2.5
  }
}

Optional parameters start and end must be provided together (in seconds). Returned timestamps remain relative to the source file start, duration is the full file duration, and usage.seconds is the processed window duration. The desktop gateway also supports passing text_path; the Headless service requires text. For advanced configuration, see normalization, dictionaries, protected terms and track/channel options.

The model parameter specifies the alignment model, which determines the supported language range. Omitting it defaults to EdgeSpeak/Lattice-2, which covers all languages in the language catalog, mixed-language text, and und. EdgeSpeak/Lattice-1 covers Chinese, English, and German only, and is used only when explicitly passing model=EdgeSpeak/Lattice-1; passing any other language or und on it will be rejected. Both models can be queried in GET /v1/models, their supported_endpoints include /v1/audio/alignments, and default_for indicates the current default tier.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@weather.wav -F text="今日はいい天気ですね。" \
  -F language=ja

Example output:

JSON
{
  "task": "align",
  "duration": 2.1,
  "language": "ja",
  "text": "今日はいい天気ですね。",
  "segments": [
    {
      "id": 0,
      "start": 0.3,
      "end": 1.9,
      "text": "今日はいい天気ですね。",
      "words": [
        { "word": "今日は", "start": 0.3, "end": 0.7, "score": 0.97 },
        { "word": "いい", "start": 0.7, "end": 0.9, "score": 0.95 },
        { "word": "天気", "start": 0.9, "end": 1.4, "score": 0.98 },
        { "word": "ですね。", "start": 1.4, "end": 1.9, "score": 0.94 }
      ]
    }
  ],
  "usage": { "type": "duration", "seconds": 2.1 }
}

Explicitly specifying model=EdgeSpeak/Lattice-1 and passing language=ja returns HTTP 422 language_unsupported. When encountering this error, switch the model rather than changing the language code.

language is a language hint for the tokenizer, not a hard capability gate. Passing codes like zh, yue, or ja explicitly specifies the language. Omitting it lets the runtime infer the language from the reference text; if it cannot, the request aligns without a language instead of failing. Text alone can confuse languages that share a writing system, such as Cantonese and Mandarin written in Chinese characters, so send the code when you know it. Passing und skips language detection and makes no language assumptions, which is suitable for mixed-language text, romanized text, and recordings without a matching language code. und does not fall back entirely to G2P: the tokenizer matches each word against dictionaries across all languages (for example, English words in mixed Chinese-English text still match English dictionary entries; only words absent from all dictionaries fall back to the generic phoneme set). Text normalization selects readings segment by segment based on text context. For example, 我想要$100, I want $100. is processed as 我想要一百美元, I want one hundred dollars.. Codes outside the language catalog (such as nan, hak, or wuu) are also supported on models with a G2P fallback; they keep their script and region subtags (for example nan-Hant), but lack dictionary pronunciations. Codes missing from dictionaries perform worse than und.

The text normalization rule set covers a subset of the language catalog. For codes outside the catalog, as well as listed languages that lack normalization rules (such as Cantonese yue), default requests automatically skip text normalization to prevent the entire alignment from failing. However, explicitly specifying top_k, classes, pins, apply_dictionary, or dictionary_ids returns HTTP 400 rather than silently dropping your configuration. The default EdgeSpeak/Lattice-2 supports all of these inputs. An explicitly selected EdgeSpeak/Lattice-1 has no G2P fallback: it rejects und and out-of-catalog codes with HTTP 400 bad_request, and rejects uncovered languages with HTTP 422 language_unsupported.

Read alignment progress

Add the progress: 1 header to the request. The response format is NDJSON; parse each line independently as JSON rather than treating it as SSE.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -N -H "progress: 1" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@hello.wav -F text="Hello world." \
  -F 'text_normalization={"enabled":false}'

Example output:

{  "progress": 0.5}
{  "progress": 1.0}
{"task":"align","duration":2.5,"text":"Hello world.","segments":[{"id":0,"start":0.2,"end":1.4,"text":"Hello world.","words":[{"word":"Hello","start":0.2,"end":0.6,"score":0.98},{"word":"world.","start":0.7,"end":1.4,"score":0.96}]}],"usage":{"type":"duration","seconds":2.5}}

The line containing "task": "align" is the valid final result. Receiving progress 1.0 alone does not guarantee alignment success. The stream ends in one of three ways:

  • a "task": "align" line: the alignment succeeded;
  • an {"error": {...}} line: the alignment failed; it has the same shape as the 422 response body, while the HTTP status of the stream stays 200;
  • neither line before the stream closes: treat it as a failure.

With search_effort=auto, a {"stage":"extended_search"} line marks the start of the wider search, and progress restarts from 0 after it.

Concurrent requests and waiting

The Headless service accepts concurrent alignment requests and waits for scheduling when upload or execution slots are full, instead of failing because slots are busy. For slot options, defaults, and scope, see CLI concurrency options.

Request and response formats are the same as above. Increase client timeouts based on audio duration and queue depth. For example, budget 30 minutes:

Shell
curl --fail-with-body --max-time 1800 "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@hello.wav -F text="Hello world."

Once completed, the task returns the alignment result containing task: "align", segments, and usage. When configuring client timeouts, account for all three phases: queueing, model loading, and inference; progress mode does not emit progress lines while waiting for an execution slot.

Handle alignment failures

When the local engine cannot align the text to the audio, the request returns HTTP 422 with an error.alignment object instead of a result, and the request is not counted in usage. Branch on error.code and error.alignment; message is an English diagnostic and may change.

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@song.mp3 -F text="$(cat lyrics.txt)"

Example output (HTTP 422; values are illustrative):

JSON
{
  "error": {
    "message": "Alignment could not be completed; the search stopped before the end of the text.",
    "type": "invalid_request_error",
    "code": "alignment_failed",
    "alignment": {
      "reason": "search_incomplete",
      "search_effort_used": "standard",
      "search_path": "whole_audio",
      "search_limit": "none",
      "retry_recommended": true,
      "estimated_extra_bytes": 3221225472
    }
  }
}
CodeMeaning
alignment_failedAlignment ran but could not be completed. reason gives the detail.
alignment_search_budget_exceededA wider search would need more memory than this computer can spare, so it did not run. Try a shorter clip, for example with start / end.
FieldDescription
reasonno_path or no_words: nothing in the audio matched the text; collapsed: the words were squeezed into implausibly short timings; search_incomplete: the search stopped before the end of the text, possibly because of music or long silence.
search_effort_usedThe search that ran: standard or extended.
search_pathwhole_audio when the whole clip was searched at once; streaming when long audio was aligned piece by piece.
search_limitnone, or the limit the search hit: budget (memory budget), graph_size (the search space is too large for a wider search), estimate_overflow (the memory estimate could not be computed), streaming (piece-by-piece alignment has no wider search).
retry_recommendedAlways present. true only for alignment_failed with reason: "search_incomplete" after a standard search, when no limit other than none was hit and the wider search fits in this computer's memory budget.
estimated_extra_bytesExpected extra memory for a wider search. It covers the search only and is an estimate, not a cap.

Fields other than retry_recommended are omitted when the service does not know them. When reason is not search_incomplete, check that the audio and the text cover the same content; a wider search will not help.

A successful alignment with no aligned words (for example, text with nothing to align) returns HTTP 200 without segments; read a missing segments as an empty list.

Retry with a wider search range

The multipart field search_effort chooses the search range:

ValueBehavior
standardDefault, also used for an empty value. One search with the standard range.
extendedOne search with a wider range. Slower and uses more memory; estimated_extra_bytes from the failed attempt tells you roughly how much more.
autoRuns standard; only if it fails with retry_recommended: true, runs extended once. Any other failure is returned as is.

Any other value returns 400 bad_request with param: "search_effort". When error.alignment.retry_recommended is true, send the same request again with search_effort=extended:

Request:

Shell
curl --fail-with-body "$EDGESPEAK_BASE_URL/audio/alignments" \
  -F model=EdgeSpeak/Lattice-2 \
  -F file=@song.mp3 -F text="$(cat lyrics.txt)" \
  -F search_effort=extended

Example output (response excerpt):

JSON
{
  "task": "align",
  "duration": 214.6,
  "text": "...",
  "segments": [{ "id": 0, "start": 12.1, "end": 16.8, "text": "...", "words": [{ "word": "...", "start": 12.1, "end": 12.5, "score": 0.93 }] }],
  "search_effort_used": "extended",
  "usage": { "type": "duration", "seconds": 214.6 }
}

Success responses include search_effort_used when the service reports the search it ran. With search_effort=auto, search_retried: true appears only when the second search actually ran; it is omitted otherwise. A failed extended search returns retry_recommended: false.

Continue: Transcription and speakers · API index