Forced alignment

You have the text. Get the timing.

Give EdgeSpeak the audio and the matching text, and it finds when each word is spoken.

Direct answer

Forced alignment locates known text in audio instead of recognizing unknown speech. EdgeSpeak returns word and segment timestamps on your computer.

Use it on scripts, lyrics and audiobooks

When you already have the text, alignment is more direct than transcribing again: same words, correct timing.

From the app, the CLI or the API

Use the Alignment page, run the command, or call the local gateway.

edgespeak-cli align meeting.m4a --text-file transcript.txt -o aligned.json

Language handling

Send the language code when you know it: Cantonese written in Chinese characters, for example, can be read as Mandarin if the language is inferred. Use und for mixed-language text, which is then matched word by word.

Feed the result to subtitles

Aligned output drives SRT, VTT and karaoke ASS export.

FAQ

How is this different from transcription?

Transcription recognizes what was said. Alignment takes the text you give it and finds where each word occurs.

Does text normalization change my spelling?

It is on by default. Turn it off to preserve the original spelling exactly.

Is there an API?

Yes: POST /v1/audio/alignments on the local gateway, with file and text.

Evidence links