Cerence vs Speechmatics

Cerence — Application · Public · $384.9M mkt cap · 4 of 4 figures sourced  |  Speechmatics — Infrastructure · Private · 1 of 1 figure sourced

Relationship

Cerence Voice Lab and Voice Agent API do comparable work on speech to text and text to speech; both also serve buyers who need to create synthetic voice and audio and turn speech into text; Speechmatics's scale not recorded.

Assembled from the recorded fields for this pair, not hand-checked. The comparison below is read from each company’s own profile.

3 of 5 capabilitiesShared product type

3 of 5 capabilities — Shares speech to text, text to speech and voice agent.

Ludbee capability tags · from the product records

Shared product type — Both ship developer tool.

Ludbee product records

Aligned comparison

FieldCerenceSpeechmatics
Size$384.9M mkt capnot disclosed
Employees1,300—
Founded20192006 13 yrs earlier
StatusPublicPrivate
CategoryApplicationInfrastructure
Stack layerAI agent, Application, Developer tool, PlatformAPI service, Developer tool
HeadquartersBurlington, USACambridge, United Kingdom

Capability overlap

Shared · 3

Speech to textText to speechVoice agent

Not verified for Speechmatics · 2

Agent orchestrationEvaluation and observability

Recorded for Cerence. Speechmatics’s product records say nothing either way — a missing record is not a missing capability.

Not verified for Cerence · 3

Data analysisSummarizationTranslation

Recorded for Speechmatics. Cerence’s product records say nothing either way — a missing record is not a missing capability.

Products, side by side

Algorithmic pairing — assembled from recorded fields, not hand-checked

Cerence

Developer tool

Cerence Voice LabDeveloper tool

Self-service cloud portal where developers upload their own audio to benchmark Cerence's speech recognition engines on word error rate, latency and confidence, and audition and tune the text-to-speech voice library in the browser.

Speechmatics

Developer tool

On-Device Speech to TextDeveloper tool

A locally-executing speech-to-text engine for Mac and Windows laptops that runs on about one CPU core plus the device's neural engine or GPU and roughly 800MB of memory, sending no audio over a network and claiming accuracy within 5% of the cloud API.

No counterpart

Cerence sells these in a stack layer with no product recorded for Speechmatics yet — nothing on the other side to compare them against.

Application

Car KnowledgeApplication

Delivers vehicle-specific insights through AI-refined search, answering questions about the vehicle.

Cerence AssistantApplication

Turnkey in-car voice assistant that carmakers deploy as-is, running core functions onboard the vehicle while reaching the cloud for live information such as news, weather and flight updates.

Chat ProApplication

Brings friendly, free-flowing small talk to the car -- contextual in-car conversation distinct from vehicle-specific Q&A.

Emergency Vehicle DetectionApplication

Acoustically detects a wide range of global sirens to alert the driver.

In-Car CommunicationApplication

Optimises conversations between vehicle occupants across rows.

Mobile Work AgentApplication

Turns drive time into productive time: a voice-based work assistant for drivers.

Ownership Companion AgentApplication

Always-available personal OEM representative that helps drivers understand and maintain their vehicle.

Speech Signal EnhancementApplication

Removes noise, unwanted sounds and speech interference from in-vehicle audio.

AI agent

Dealer Assist AgentAI agent

Handles dealership customer inquiries as part of Cerence's automotive AI agents portfolio.

Platform

Cerence xUIPlatform

Hybrid generative-AI platform for the car cabin built on Cerence's own CaLLM model family, splitting work between the vehicle's embedded hardware and the cloud so the assistant still answers with no connection.

Speechmatics sells these in a stack layer with no product recorded for Cerence yet — nothing on the other side to compare them against.

API service

Speech to Text APIAPI service

Speechmatics' automatic speech recognition API transcribes audio into text in 55+ languages in either real-time streaming or batch mode, with speaker diarization, custom dictionary, translation and summarization options.

SummarizationAPI service

Generates a short summary of an audio file in the same API call that transcribes it, as paragraphs or bullets. It is a Speech Intelligence feature enabled by adding a config block to a batch Speech to Text job, not a product bought on its own.

Text to SpeechAPI service

Speechmatics' text-to-speech API generates streaming synthetic English speech from text with sub-150ms latency using four named voices (Sarah, Theo, Megan, Jack), aimed at real-time voice agent use.

TranslationAPI service

Translates a transcript into other languages in the same API call that produces it, for files or live audio. It is a feature switched on inside a Speech to Text request rather than a separate product; the docs file it under Speech to Text and return the translations alongside the transcript.

Voice Agent APIAPI service

Speechmatics' voice agent offering provides a real-time conversational speech API - including the Flow WebSocket endpoint that chains speech-to-text, an LLM, text-to-speech and function calling - plus a Python Voice SDK for turn detection and speaker management, and integrations with Vapi, LiveKit and Pipecat.