How is this different from transcription?
Transcription recognizes what was said. Alignment takes the text you give it and finds where each word occurs.
Forced alignment
Give EdgeSpeak the audio and the matching text, and it finds when each word is spoken.
Forced alignment locates known text in audio instead of recognizing unknown speech. EdgeSpeak returns word and segment timestamps on your computer.
When you already have the text, alignment is more direct than transcribing again: same words, correct timing.
Use the Alignment page, run the command, or call the local gateway.
edgespeak-cli align meeting.m4a --text-file transcript.txt -o aligned.json
Send the language code when you know it: Cantonese written in Chinese characters, for example, can be read as Mandarin if the language is inferred. Use und for mixed-language text, which is then matched word by word.
Aligned output drives SRT, VTT and karaoke ASS export.
Transcription recognizes what was said. Alignment takes the text you give it and finds where each word occurs.
It is on by default. Turn it off to preserve the original spelling exactly.
Yes: POST /v1/audio/alignments on the local gateway, with file and text.