Speakers

See who spoke when, on your computer

Turn on speaker detection when transcribing and segments get a speaker label where one can be assigned.

Direct answer

EdgeSpeak labels speakers per segment on your computer. Rename them, save speaker profiles, and match the same voice across recordings.

Name the speakers

Rename speaker labels and save profiles. Voice profiles are encrypted and stay on your computer.

Use it from scripts and agents

transcribe --diarize in the CLI, diarized_json from the local gateway, /v1/speaker/diarizations for a speaker timeline, and the edgespeak_diarize_file MCP tool.

edgespeak-cli transcribe input.m4a --diarize -o out.json

Keep speakers in exports

JSON exports keep each segment's speaker label, which is empty when no speaker was assigned; karaoke subtitles can prefix speaker names.

FAQ

Are speaker labels real identities?

No. They are voice groups within the recording until you name them.

Does it need an extra download?

Yes, the optional speaker model (about 206 MB), so plain transcription stays small.

Can I tell it how many speakers there are?

Yes. The API and MCP take an optional num_speakers from 1 to 32 as a hint for grouping voices.

Evidence links