เอนจินเสียงบนอุปกรณ์

นำโมเดล AI เสียงขนาดใหญ่ไว้ในคอมพิวเตอร์ของคุณ

EdgeSpeak บีบอัดโมเดลเสียงระดับมืออาชีพสำหรับเดสก์ท็อป และถอดเสียงการประชุม สัมภาษณ์ วิดีโอ และไฟล์บันทึกในเครื่อง ไฟล์เสียง วิดีโอ และข้อความยังอยู่บนอุปกรณ์ของคุณ

  • ถอดเสียง/วิดีโอในเครื่องได้แล้ว
  • Gateway ในเครื่องและ CLI ใช้ได้แล้ว
  • อัปเดตโมเดลและความสามารถเสียงในอนาคต
ตัวอย่าง transcript
ตัวอย่าง transcriptlocal-transcribe.mov

นักพัฒนา · CLI & Skills

ติดตั้ง CLI ด้วยคำสั่งเดียว

ติดตั้งด้วยบรรทัดเดียวแบบที่คุ้นเคย เหมือนเครื่องมือนักพัฒนาที่คุณใช้อยู่แล้ว CLI ทำงานได้ในตัวเอง ไม่ต้องใช้แอปเดสก์ท็อป และถอดเสียงได้จากเทอร์มินัลโดยตรง

ติดตั้งผ่าน Terminal บน macOS หรือ Linux และผ่าน PowerShell บน Windows x64 โดย CPU ใช้งานได้ทันที Linux จะตรวจพบ NVIDIA CUDA อัตโนมัติ ส่วน Windows ให้เลือก CUDA เมื่อต้องการ

macOS · Apple SiliconWindows x64 · CPU / CUDALinux x86_64 · CPU / CUDA
อ่านเอกสาร CLI

macOS / Linux · Terminal

curl -fsSL https://edgespeak.com/install.sh | sh

Windows · PowerShell

irm https://edgespeak.com/install.ps1 | iex
edgespeak-cli login
edgespeak-cli transcribe media.mp4 -o media.srt
edgespeak-cli generate "Summarize this transcript" --model Qwen/Qwen3.5-4B

Agent Skills เครื่องยนต์เดียวกัน

EdgeSpeak Skills สอนให้ Claude Code, Cursor และเอเจนต์ที่รองรับ Skills ถอดเสียงบนอุปกรณ์ผ่าน CLI โอเพนซอร์สบน GitHub

npx skills add lattifai/EdgeSpeak

ถอดเสียงเร็วบนอุปกรณ์

นำเข้าเสียงหรือวิดีโอ แล้วถอดเสียงในเครื่องด้วยความแม่นยำสูง

ใส่การประชุม สัมภาษณ์ หรือวิดีโอเข้ามา EdgeSpeak จะทำความเข้าใจเสียง สร้างข้อความถอดเสียง และจัดแนวข้อความบนคอมพิวเตอร์นี้ เมื่อเครื่องมือต่อไปต้องทำงานต่อ ก็รับข้อความจาก gateway ในเครื่องได้

  1. EdgeSpeak บนอุปกรณ์ของคุณ

    EdgeSpeak เปลี่ยน audio/video เป็นข้อความบนคอมพิวเตอร์ตรงหน้าคุณ

  2. แก้ไปพร้อมกับฟัง

    ข้อความตามการเล่นเสียง ทำให้ตรวจและแก้ได้ในพื้นที่เดียว

  3. ผลลัพธ์ไปต่อได้

    ส่งออกข้อความหรือซับไตเติล หรือให้เครื่องมืออื่นใช้ transcript ที่เสร็จแล้วต่อ

โมเดลเฉพาะบนอุปกรณ์

Lattice-2 สร้างมาเพื่อเดสก์ท็อป

Lattice-2 ผ่านการบีบอัดและปรับ inference บนอุปกรณ์ ให้คอมพิวเตอร์ทั่วไปประมวลผลการประชุม สัมภาษณ์ วิดีโอ และไฟล์บันทึกได้อย่างมีประสิทธิภาพ

การบีบอัดและ inference บนอุปกรณ์

โมเดล AI เสียงถูกปรับให้เหมาะกับการทำงานบนเดสก์ท็อป ลดเวลารอ และการถอดเสียงในเครื่องไม่ต้องพึ่งบริการภายนอก

โมเดลคู่ Flash / Pro

Flash ตอบสนองเร็วสำหรับการประชุมและวิดีโอทั่วไป ส่วน Pro เหมาะกับเสียงที่ยากกว่า สำเนียงซับซ้อนกว่า และต้องการความแม่นยำสูงกว่า โดยใช้ทรัพยากรในเครื่องมากขึ้น

เชื่อมเข้ากับ workflow ของนักพัฒนา

Lattice-2 รองรับมากกว่า 40 ภาษา หลายสำเนียงอังกฤษ และภาษาถิ่นจีน เอนจินในเครื่องเดียวกันส่งต่อให้ CLI, agent และ automation ผ่าน gateway ได้

LOCAL LM + VLM · 0.6B — 27B

ภาษาและภาพแบบภายในเครื่องในรันไทม์ EdgeSpeak เดียวกัน

EdgeSpeak มีโมเดล Qwen และ Gemma ตั้งแต่ 0.6B ถึง 27B โหลดครั้งละหนึ่งโมเดลและพักอัตโนมัติเมื่อไม่ได้ใช้งาน

สร้างด้วยโมเดลภายในเครื่อง
Qwen/Qwen3-0.6B0.6B0.4 GB
Qwen/Qwen3.5-0.8B0.8B0.5 GB
Qwen/Qwen3.5-2B2B1.3 GB
google/gemma-4-E2B-it2B3.1 GB
Qwen/Qwen3.5-4Bค่าเริ่มต้น4B2.7 GB
google/gemma-4-E4B-it4B5.0 GB
Qwen/Qwen3.5-9B9B5.7 GB
google/gemma-4-12B-it12B7.1 GB
Qwen/Qwen3.6-27B27B16.8 GB

ข้อความ + รูปภาพ · ใช้ร่วมกับ OpenAI · ภายในเครื่องก่อน

แอปเดสก์ท็อปจริง

EdgeSpeak ทำงานบนคอมพิวเตอร์ของคุณโดยตรง

นำเข้าเสียงหรือวิดีโอ เลือกโมเดลในเครื่อง แล้วส่งออกผลลัพธ์ เมื่อต้องการ automation ให้ส่งงานต่อไปยัง CLI หรือ agent ผ่าน gateway ในเครื่อง

เสียง วิดีโอ และข้อความถอดเสียงใน workspace เดียวการเล่น ไทม์ไลน์ ข้อความถอดเสียง export ไฟล์ล่าสุด และโมเดล EdgeSpeak ที่ใช้งานอยู่รวมกัน ลดการสลับเครื่องมือและรักษาบริบทไว้
หน้าจอ Transcribe ของ EdgeSpeak desktop แสดง workspace สถานะโมเดล EdgeSpeak ในเครื่อง ข้อความถอดเสียง ตัวควบคุมการเล่น และไฟล์ล่าสุด
เลือกโมเดล EdgeSpeak ให้เหมาะกับงานEdgeSpeak รองรับมากกว่า 40 ภาษา หลายสำเนียงอังกฤษ และภาษาถิ่นจีน Flash เร็วและแข็งแรง ส่วน Pro แม่นยำกว่าและใช้ทรัพยากรในเครื่องมากกว่า
หน้าจอ Models ของ EdgeSpeak desktop แสดงตัวเลือก EdgeSpeak Flash และ EdgeSpeak Pro ในเครื่อง
ให้เครื่องมืออื่นใช้ EdgeSpeak ได้CLI, agent และ automation ส่งเสียงให้ EdgeSpeak แล้วรับข้อความถอดเสียงกลับจาก EdgeSpeak ได้
หน้าจอ Gateway ของ EdgeSpeak desktop แสดงวิธีที่เครื่องมือในคอมพิวเตอร์เดียวกันใช้เอนจินเสียงในเครื่อง

Workflows

เริ่มจาก workflow เสียงของคุณ

สำรวจไฟล์ส่วนตัว การถอดเสียงและวิดีโอในเครื่อง agent automation และตัวเลือกในเครื่องเทียบกับ cloud

ความเป็นส่วนตัว

ประมวลผลเสียงแบบ local-first.

EdgeSpeak is designed so imported media and generated transcripts stay on your device unless you export, upload, or share them.

การบีบอัดและ inference บนอุปกรณ์

โมเดล AI เสียงถูกปรับให้เหมาะกับการทำงานบนเดสก์ท็อป ลดเวลารอ และการถอดเสียงในเครื่องไม่ต้องพึ่งบริการภายนอก

ผลลัพธ์ไปต่อได้

ส่งออกข้อความหรือซับไตเติล หรือให้เครื่องมืออื่นใช้ transcript ที่เสร็จแล้วต่อ

Speech gateway ในเครื่อง

ให้ agent เรียก EdgeSpeak ได้โดยตรง

EdgeSpeak มี API เสียงในเครื่องที่เข้ากันได้กับ OpenAI ให้ CLI, agent และ automation ถอดเสียงบนคอมพิวเตอร์เครื่องเดียวกัน ไม่ใช่ cloud อีกแห่ง แต่เป็น gateway เสียงของคอมพิวเตอร์คุณ

Endpoint ในเครื่องที่เข้ากันได้กับ OpenAICLI / agent / automationEdgeSpeak Flash และ Pro
POST /v1/audio/transcriptionslocalhost:1117
curl http://127.0.0.1:1117/v1/audio/transcriptions \
  -H "Authorization: Bearer sk-edgespeak-..." \
  -F file=@meeting.m4a \
  -F model="lattice-2-flash"

Choose by workflow

Local software, cloud APIs, or a model you maintain

The right option depends on where media can go, how much workflow you want ready-made, and who should maintain the speech stack.

Choose by workflowEdgeSpeakCloud transcriptionSelf-managed local model
Source mediaStays on this deviceUploaded to a remote serviceStays on the machine you configure
Review workflowDesktop playback, transcript, and exportDepends on the providerYou build the review surface
Tools and agentsLocal CLI and OpenAI-compatible gatewayHosted APIYou build and maintain the integration
OperationsInstall the app and local modelsManage an account, API keys, and network accessMaintain the runtime, models, and dependencies
Read the full comparison

Early bird

ใช้ EdgeSpeak ตลอดชีพในราคา $49

เวอร์ชันปัจจุบันมีการถอดเสียงและวิดีโอในเครื่อง gateway ในเครื่อง และ CLI ซื้อครั้งเดียวและรับอัปเดตโมเดลกับความสามารถด้านเสียงในอนาคต

ดาวน์โหลด

ติดตั้ง EdgeSpeak บน macOS.

Choose the current macOS build or open the Windows Beta page for the x64 installer and verified setup instructions.

Current macOS build

Use the available desktop build today. Your account keeps purchase, license, and device management in one place.

Windows build

Download the Windows x64 Beta, verify its SHA-256, and follow the installation guide. This Beta uses manual updates.

License and devices

Sign in with the email used for purchase, then manage your license key and activated devices from your account.

FAQ

Frequently asked questions

On-device processing, supported platforms, models, and the lifetime license — answered before you buy.

See all questions
01Does EdgeSpeak upload source media?

Imported media and generated transcripts stay on your device unless you export, upload, or share them yourself.

02Which desktop platforms are available?

EdgeSpeak is available for Apple Silicon Macs with macOS 14.0 or later. A Windows 10/11 x64 Beta is also available for manual installation; Windows in-app updates are not enabled yet.

03What does the lifetime license include?

The early-bird lifetime license is a one-time purchase for permanent use, future model and speech-capability updates, and up to four activated devices.

04What is the difference between Flash and Pro?

Flash is tuned for faster everyday transcription. Pro raises the accuracy ceiling for harder audio and uses more local resources.

05Can CLI tools and AI agents use EdgeSpeak?

Yes. The bundled CLI and local OpenAI-compatible gateway let trusted tools on the same computer call the local speech engine.

เสียงจากชุมชน

พัฒนา EdgeSpeak ด้วย workflow จริง

แชร์ workflow จริง ความต้องการเชื่อมต่อ agent และไอเดียผลิตภัณฑ์บน Discord

  • Workflow จริง
  • Agent / API
  • ไอเดียผลิตภัณฑ์

ความคิดเห็นของคุณเข้าสู่การพัฒนาผลิตภัณฑ์โดยตรง

เอนจินเสียงบนอุปกรณ์

นำโมเดล AI เสียงขนาดใหญ่ไว้ในคอมพิวเตอร์ของคุณ

Transcribe locally, review against an accurate timeline, then export or continue through the local gateway.