Furhat Robotics vs Speechmatics

Furhat Robotics — Application · Private · 1 of 1 figure sourced  |  Speechmatics — Infrastructure · Private · 1 of 1 figure sourced

Relationship

Furhat robot and Voice Agent API do comparable work on speech to text, text to speech and voice agent; both also serve buyers who need to create synthetic voice and audio; scale not recorded for either; ships API service and developer tool rather than the same layer.

Assembled from the recorded fields for this pair, not hand-checked. The comparison below is read from each company’s own profile.

3 of 4 capabilitiesDifferent layer

3 of 4 capabilities — Shares speech to text, text to speech and voice agent.

Ludbee capability tags · from the product records

Different layer — Speechmatics ships API service and developer tool, not the same layer.

Ludbee product records

Aligned comparison

FieldFurhat RoboticsSpeechmatics
Sizenot disclosednot disclosed
Employees——
Founded20142006 8 yrs earlier
StatusPrivatePrivate match
CategoryApplicationInfrastructure
Stack layerPlatform, RobotAPI service, Developer tool
HeadquartersStockholm, SwedenCambridge, United Kingdom

Capability overlap

Shared · 3

Speech to textText to speechVoice agent

Not verified for Speechmatics · 1

Text generation

Recorded for Furhat Robotics. Speechmatics’s product records say nothing either way — a missing record is not a missing capability.

Not verified for Furhat Robotics · 3

Data analysisSummarizationTranslation

Recorded for Speechmatics. Furhat Robotics’s product records say nothing either way — a missing record is not a missing capability.

Products, side by side

Algorithmic pairing — assembled from recorded fields, not hand-checked

Furhat Robotics

No shared stack layer with the other side.

Speechmatics

No shared stack layer with the other side.

No counterpart

Furhat Robotics sells these in a stack layer with no product recorded for Speechmatics yet — nothing on the other side to compare them against.

Platform

FurhatAIPlatform

Software subscription that adds model-driven characters, voices and dialogue tools to the Furhat robot or SDK.

Robot

Furhat robotRobot

Desktop social robot with a back-projected animated face, used for conversational research and public-facing installations.

Speechmatics sells these in a stack layer with no product recorded for Furhat Robotics yet — nothing on the other side to compare them against.

API service

Speech to Text APIAPI service

Speechmatics' automatic speech recognition API transcribes audio into text in 55+ languages in either real-time streaming or batch mode, with speaker diarization, custom dictionary, translation and summarization options.

SummarizationAPI service

Generates a short summary of an audio file in the same API call that transcribes it, as paragraphs or bullets. It is a Speech Intelligence feature enabled by adding a config block to a batch Speech to Text job, not a product bought on its own.

Text to SpeechAPI service

Speechmatics' text-to-speech API generates streaming synthetic English speech from text with sub-150ms latency using four named voices (Sarah, Theo, Megan, Jack), aimed at real-time voice agent use.

TranslationAPI service

Translates a transcript into other languages in the same API call that produces it, for files or live audio. It is a feature switched on inside a Speech to Text request rather than a separate product; the docs file it under Speech to Text and return the translations alongside the transcript.

Voice Agent APIAPI service

Speechmatics' voice agent offering provides a real-time conversational speech API - including the Flow WebSocket endpoint that chains speech-to-text, an LLM, text-to-speech and function calling - plus a Python Voice SDK for turn detection and speaker management, and integrations with Vapi, LiveKit and Pipecat.

Developer tool

On-Device Speech to TextDeveloper tool

A locally-executing speech-to-text engine for Mac and Windows laptops that runs on about one CPU core plus the device's neural engine or GPU and roughly 800MB of memory, sending no audio over a network and claiming accuracy within 5% of the cloud API.