EdgeSpeak auf deinem Gerät
EdgeSpeak macht Audio und Video auf deinem Computer zu Text.
On-device Speech Engine
EdgeSpeak komprimiert professionelle Sprachmodelle für den Desktop und transkribiert Meetings, Interviews, Videos und Aufnahmen lokal. Audio, Video und Transkripte bleiben auf deinem Gerät.

Für Entwickler · CLI & Skills
Der vertraute Einzeiler, den du von deinen Entwicklertools kennst. Das CLI ist eigenständig, braucht keine Desktop-App und transkribiert direkt im Terminal.
Installation unter macOS und Linux im Terminal, unter Windows x64 in PowerShell. CPU funktioniert sofort; Linux erkennt NVIDIA CUDA automatisch, unter Windows ist CUDA optional auswählbar.
macOS / Linux · Terminal
curl -fsSL https://edgespeak.com/install.sh | sh
Windows · PowerShell
irm https://edgespeak.com/install.ps1 | iex
EdgeSpeak Skills bringen Claude Code, Cursor und anderen Skills-fähigen Agents bei, über das CLI direkt auf dem Gerät zu transkribieren. Open Source auf GitHub.
npx skills add lattifai/EdgeSpeak
Schnelle On-device-Transkription
Ziehe ein Meeting, Interview oder Video hinein. EdgeSpeak übernimmt Sprachverstehen, Transkripterzeugung und Textausrichtung auf diesem Computer; nachgelagerte Tools holen das Transkript über das lokale Gateway ab.
EdgeSpeak macht Audio und Video auf deinem Computer zu Text.
Das Transkript folgt der Wiedergabe, damit Korrekturen in einem Arbeitsbereich bleiben.
Exportiere Text oder Untertitel, oder lass ein anderes Tool mit dem fertigen Transkript weiterarbeiten.
Dediziertes On-device-Modell
Lattice-2 ist komprimiert und für On-Device-Inferenz optimiert, damit ein normaler Computer Meetings, Interviews, Videos und Aufnahmen effizient verarbeitet.
Das Sprach-KI-Modell wird für Desktop-Ausführung geformt, reduziert Wartezeit und lokale Transkription hängt nicht von externen Diensten ab.
Flash ist für alltägliche Meetings und Videotranskription schnell. Pro zielt auf schwierigere Audios, komplexere Akzente und höhere Genauigkeit, nutzt dabei mehr lokale Ressourcen.
Lattice-2 unterstützt über 40 Sprachen, mehrere englische Akzente und chinesische Dialekte; dieselbe lokale Engine steht CLI, Agents und Automatisierung über das Gateway bereit.
LOCAL LM + VLM · 0.6B — 27B
EdgeSpeak bietet jetzt Qwen- und Gemma-Modelle von 0.6B bis 27B. Es wird jeweils ein Modell geladen und bei Leerlauf pausiert.
Mit einem lokalen Modell generierenQwen/Qwen3-0.6B0.6B0.4 GBQwen/Qwen3.5-0.8B0.8B0.5 GBQwen/Qwen3.5-2B2B1.3 GBgoogle/gemma-4-E2B-it2B3.1 GBQwen/Qwen3.5-4BStandard4B2.7 GBgoogle/gemma-4-E4B-it4B5.0 GBQwen/Qwen3.5-9B9B5.7 GBgoogle/gemma-4-12B-it12B7.1 GBQwen/Qwen3.6-27B27B16.8 GBText + Bilder · OpenAI-kompatibel · lokal zuerst
Echte Desktop-App
Importiere Audio oder Video, wähle ein lokales Modell und exportiere das Ergebnis. Für Automatisierung übergibst du Aufgaben über das lokale Gateway an CLI oder Agents.



Workflows
Entdecke private Dateien, lokale Audio- und Videotranskription, Agent-Automatisierung sowie lokale und Cloud-Optionen.
Keep source media on-device and generate the transcript locally.
Review with playback, then export text or subtitles.
Install the EdgeSpeak Skill and let agents call the local CLI.
Choose by media location, processing scale, and integration path.
Datenschutz
EdgeSpeak is designed so imported media and generated transcripts stay on your device unless you export, upload, or share them.
Das Sprach-KI-Modell wird für Desktop-Ausführung geformt, reduziert Wartezeit und lokale Transkription hängt nicht von externen Diensten ab.
Exportiere Text oder Untertitel, oder lass ein anderes Tool mit dem fertigen Transkript weiterarbeiten.
Lokales Speech Gateway
EdgeSpeak bietet eine lokale, OpenAI-kompatible Sprach-API. CLI, Agents und Automatisierungen transkribieren auf demselben Computer. Keine weitere Cloud, sondern das Sprach-Gateway deines Computers.
POST /v1/audio/transcriptionslocalhost:1117curl http://127.0.0.1:1117/v1/audio/transcriptions \ -H "Authorization: Bearer sk-edgespeak-..." \ -F file=@meeting.m4a \ -F model="lattice-2-flash"
Choose by workflow
The right option depends on where media can go, how much workflow you want ready-made, and who should maintain the speech stack.
| Choose by workflow | EdgeSpeak | Cloud transcription | Self-managed local model |
|---|---|---|---|
| Source media | Stays on this device | Uploaded to a remote service | Stays on the machine you configure |
| Review workflow | Desktop playback, transcript, and export | Depends on the provider | You build the review surface |
| Tools and agents | Local CLI and OpenAI-compatible gateway | Hosted API | You build and maintain the integration |
| Operations | Install the app and local models | Manage an account, API keys, and network access | Maintain the runtime, models, and dependencies |
Preise
Die aktuelle Version umfasst lokale Audio- und Videotranskription, das lokale Gateway und CLI. Einmal kaufen und künftige Modell- und Sprachfunktionsupdates erhalten.
Aktuell $49, regulär $99. Lifetime Access, einmaliger Kauf während Early Bird.
Mehr Geräte oder Teams: sales@edgespeak.com
Download
Lade den aktuellen Apple-Silicon-Build, prüfe die Release-Metadaten und aktiviere deine license in Sekunden.
Use the available desktop build today. Your account keeps purchase, license, and device management in one place.
Download the Windows x64 Beta, verify its SHA-256, and follow the installation guide. This Beta uses manual updates.
Sign in with the email used for purchase, then manage your license key and activated devices from your account.
FAQ
On-device processing, supported platforms, models, and the lifetime license — answered before you buy.
See all questionsImported media and generated transcripts stay on your device unless you export, upload, or share them yourself.
EdgeSpeak is available for Apple Silicon Macs with macOS 14.0 or later. A Windows 10/11 x64 Beta is also available for manual installation; Windows in-app updates are not enabled yet.
The early-bird lifetime license is a one-time purchase for permanent use, future model and speech-capability updates, and up to four activated devices.
Flash is tuned for faster everyday transcription. Pro raises the accuracy ceiling for harder audio and uses more local resources.
Yes. The bundled CLI and local OpenAI-compatible gateway let trusted tools on the same computer call the local speech engine.
Community-Feedback
Teile reale Workflows, Anforderungen an Agent-Integrationen und Produktideen auf Discord.
Dein Feedback fließt direkt in das Produkt ein.
On-device Speech Engine
Transcribe locally, review against an accurate timeline, then export or continue through the local gateway.