Speech to text

Turn speech into text without uploading it

Import a recording or a video, and EdgeSpeak transcribes it on your computer while you review it against playback.

Direct answer

EdgeSpeak is a desktop app that transcribes audio and video locally with Lattice-2 Flash or Pro, which support more than 40 languages. Files stay on your machine.

Review while you listen

The transcript follows playback. Click a line to jump to it, fix a word in place, and export when it reads right.

Timestamps and structure

Get word-level timestamps, sentence segmentation and optional speaker labels. Leave the language empty and it is detected from the audio.

Export or automate

Export plain text, SRT, VTT or JSON. The same engine runs from the command line and the OpenAI-compatible local gateway.

edgespeak-cli transcribe meeting.m4a -o meeting.json

Works on video too

Drop in an MP4 or another video file and get the spoken text with timing, ready to become subtitles.

FAQ

Does the audio leave my computer?

Not when you use the on-device engine. Transcription runs locally; your recordings are not uploaded to transcribe them.

Which languages are supported?

Lattice-2 Flash and Pro support more than 40 languages. The language guide lists the codes and what each capability supports.

Can I get timestamps for every word?

Yes. Word-level timestamps are available in the JSON export, the CLI and the local gateway.

Can I call it from my own scripts?

Yes. Use edgespeak-cli, or send audio to the OpenAI-compatible local gateway on your computer.

Evidence links