Features
Everything EdgeSpeak does with speech, on your computer
Transcription, subtitles, chat, speech and timing in one desktop app. Pick a capability to see how it works.
Speech to text
Speech to Text on Your Computer
Turn audio and video into text on your own computer. 40+ languages, word-level timestamps, and export to text, SRT, VTT or JSON.
Custom Dictionary for Names and Terms
Add names, product terms and jargon to the EdgeSpeak dictionary so they are recognized correctly in transcription and translated your way.
Subtitles
Subtitle Generator: SRT, VTT and ASS
Generate subtitles from audio or video on your computer and export SRT, VTT or ASS with accurate word-level timing.
Karaoke Subtitles with Word-by-Word Highlight
Export ASS karaoke subtitles where each word lights up as it is spoken. Choose a highlight preset, fonts, colors and canvas size.
Forced Alignment: Word Timestamps for Known Text
Align a script or transcript you already have to the audio and get word and segment timestamps, from the app, CLI or API.
Voice and chat
Voice and Text Chat with Local AI Models
Talk to an AI model by voice or text on your computer. Pick a persona such as Buddy, or switch to live interpretation between two languages.
Text to Speech on Your Computer
Generate natural speech from text locally. Streaming playback starts as audio is produced, with reusable voices and WAV output.