On-Device Speech to Text
A locally-executing speech-to-text engine for Mac and Windows laptops that runs on about one CPU core plus the device's neural engine or GPU and roughly 800MB of memory, sending no audio over a network and claiming accuracy within 5% of the cloud API.
Find alternatives to On-Device Speech to Text
The page states verbatim: 'Pricing is based on your deployment volume and use case. Speak to our sales team for a tailored quote and volume-based discounts.' No rate appears here or on speechmatics.com/pricing, which covers only the cloud Speech-to-Text, add-ons and Text-to-Speech - the on-device offering has no published figure, so no price floor is recorded.
What it does
- Speech to text
Sources
- Pricing
-
Official documentation · 5 Sep 2026
Quote-only; verified the self-serve pricing page carries no on-device line.
- Description
-
Official documentation · 5 Sep 2026
Page title 'On-Device Speech-to-Text for Laptop'; H1 'Speech-to-Text optimized for laptop'. States macOS 14 Sonoma+ (M1 or newer) and Windows 11 (Intel/AMD/ARM, 2GB VRAM min), '~800MB of system memory', 'within 5% of cloud accuracy', full speaker diarization and ID, CoreML/DirectML optimization, and an Adobe Premiere deployment.
- Sold within
-
Live re-check 2026-09-15: the two candidate pages contradict each other — one implies a separately-licensed product, the other frames on-device as just one deployment mode of the same underlying API.
Also from Speechmatics
-
Speech to Text API API service
Speechmatics' automatic speech recognition API transcribes audio into text in 55+ languages in either real-time streaming or batch mode, with speaker diarization, custom dictionary, translation and summarization options.
-
Text to Speech API service
Speechmatics' text-to-speech API generates streaming synthetic English speech from text with sub-150ms latency using four named voices (Sarah, Theo, Megan, Jack), aimed at real-time voice agent use.
-
Voice Agent API API service
Speechmatics' voice agent offering provides a real-time conversational speech API - including the Flow WebSocket endpoint that chains speech-to-text, an LLM, text-to-speech and function calling - plus a Python Voice SDK for turn detection and speaker management, and integrations with Vapi, LiveKit and Pipecat.
-
Translation API service
Translates a transcript into other languages in the same API call that produces it, for files or live audio. It is a feature switched on inside a Speech to Text request rather than a separate product; the docs file it under Speech to Text and return the translations alongside the transcript.
-
Summarization API service
Generates a short summary of an audio file in the same API call that transcribes it, as paragraphs or bullets. It is a Speech Intelligence feature enabled by adding a config block to a batch Speech to Text job, not a product bought on its own.
Open Speechmatics in the directory
Something wrong here? Send a correction — quote this product id: speechmatics-on-device-speech-to-text.