Speech to Text API
Speechmatics' automatic speech recognition API transcribes audio into text in 55+ languages in either real-time streaming or batch mode, with speaker diarization, custom dictionary, translation and summarization options.
Find alternatives to Speech to Text API
Free plan gives $100 in credit with no credit card and 2 concurrent real-time sessions. Pro is metered per hour of audio: Batch Melia 1 $0.129/hr, Batch Standard $0.24/hr, Batch Enhanced $0.40/hr, Real-time Standard $0.24/hr, Real-time Enhanced $0.43/hr, with a 20% discount over 500 hr/month. Metered add-ons on the same job: Translation $0.65/hr, Summaries $0.12/hr, Chapters $0.40/hr, Sentiment $0.12/hr, Topics $0.20/hr. Enterprise is custom-quoted with volume discounts from 24,000 hours annually. No monthly subscription fee is published, hence a floor of 0.
What it does
- Speech to text
- Translation
- Summarization
- Data analysis
How it compares
-
Deepgram Speech-to-Text
Deepgram Alternative: Better Accuracy, 55+ Languages, Lower Cost | Speechmatics
Official documentation · 19 Sep 2026 -
AssemblyAI Universal Speech-to-Text API
Speechmatics vs AssemblyAI: Speech-to-Text Comparison (2026)
Company disclosure · 23 Sep 2026
Sources
- Pricing
- Description
-
Official documentation · 5 Sep 2026
Product page markets a Speech-to-Text API across 55+ languages, real-time and batch, sub-500ms latency, speaker diarization, alphanumeric recognition.
- Sold within
-
Live re-check 2026-09-15: speechmatics.com/pricing shows Speech-to-Text as its own line with independent price tiers, no parent packaging stated.
- Deployment
-
Official documentation · 5 Sep 2026
Lists On-Prem ('Docker containers, Kubernetes, or a preconfigured virtual appliance'), Cloud and On-Device as the deployment modes; docs.speechmatics.com/deployments documents CPU/GPU speech-to-text containers, a language-ID container and a translation container.
Also from Speechmatics
-
Text to Speech API service
Speechmatics' text-to-speech API generates streaming synthetic English speech from text with sub-150ms latency using four named voices (Sarah, Theo, Megan, Jack), aimed at real-time voice agent use.
-
Voice Agent API API service
Speechmatics' voice agent offering provides a real-time conversational speech API - including the Flow WebSocket endpoint that chains speech-to-text, an LLM, text-to-speech and function calling - plus a Python Voice SDK for turn detection and speaker management, and integrations with Vapi, LiveKit and Pipecat.
-
On-Device Speech to Text Developer tool
A locally-executing speech-to-text engine for Mac and Windows laptops that runs on about one CPU core plus the device's neural engine or GPU and roughly 800MB of memory, sending no audio over a network and claiming accuracy within 5% of the cloud API.
-
Translation API service
Translates a transcript into other languages in the same API call that produces it, for files or live audio. It is a feature switched on inside a Speech to Text request rather than a separate product; the docs file it under Speech to Text and return the translations alongside the transcript.
-
Summarization API service
Generates a short summary of an audio file in the same API call that transcribes it, as paragraphs or bullets. It is a Speech Intelligence feature enabled by adding a config block to a batch Speech to Text job, not a product bought on its own.
Open Speechmatics in the directory
Something wrong here? Send a correction — quote this product id: speechmatics-speech-to-text.