Cerence vs Hume AI

Cerence — Application · Public · $384.9M mkt cap · 4 of 4 figures sourced  |  Hume AI — Infrastructure · Private

Relationship

Cerence Voice Lab and Octave TTS do comparable work on text to speech; both also serve buyers who need to create synthetic voice and audio; Hume AI's scale not recorded.

Assembled from the recorded fields for this pair, not hand-checked. The comparison below is read from each company’s own profile.

3 of 5 capabilitiesShared product type

3 of 5 capabilities — Shares evaluation and observability, text to speech and voice agent.

Ludbee capability tags · from the product records

Shared product type — Both ship application and platform.

Ludbee product records

Aligned comparison

FieldCerenceHume AI
Size$384.9M mkt capnot disclosed
Employees1,300—
Founded2019—
StatusPublicPrivate
CategoryApplicationInfrastructure
Stack layerAI agent, Application, Developer tool, PlatformAPI service, Application, Data service, Platform
HeadquartersBurlington, USANew York, USA

Capability overlap

Shared · 3

Evaluation and observabilityText to speechVoice agent

Not verified for Hume AI · 2

Agent orchestrationSpeech to text

Recorded for Cerence. Hume AI’s product records say nothing either way — a missing record is not a missing capability.

Not verified for Cerence · 3

Audio editingData analysisData labelling

Recorded for Hume AI. Cerence’s product records say nothing either way — a missing record is not a missing capability.

Products, side by side

Algorithmic pairing — assembled from recorded fields, not hand-checked

Cerence

Application

Car KnowledgeApplication

Delivers vehicle-specific insights through AI-refined search, answering questions about the vehicle.

Cerence AssistantApplication

Turnkey in-car voice assistant that carmakers deploy as-is, running core functions onboard the vehicle while reaching the cloud for live information such as news, weather and flight updates.

Chat ProApplication

Brings friendly, free-flowing small talk to the car -- contextual in-car conversation distinct from vehicle-specific Q&A.

Emergency Vehicle DetectionApplication

Acoustically detects a wide range of global sirens to alert the driver.

In-Car CommunicationApplication

Optimises conversations between vehicle occupants across rows.

Mobile Work AgentApplication

Turns drive time into productive time: a voice-based work assistant for drivers.

Ownership Companion AgentApplication

Always-available personal OEM representative that helps drivers understand and maintain their vehicle.

Speech Signal EnhancementApplication

Removes noise, unwanted sounds and speech interference from in-vehicle audio.

Platform

Cerence xUIPlatform

Hybrid generative-AI platform for the car cabin built on Cerence's own CaLLM model family, splitting work between the vehicle's embedded hardware and the cloud so the assistant still answers with no connection.

Hume AI

Application

Creator StudioApplication

Creator Studio is Hume's web application, hosted at app.hume.ai, for turning written documents into narrated audio: 'From document to audio, effortlessly. Import your content, assign voices, add acting directions, and export production-ready audio.' Hume targets it at creators producing audiobooks, videos, podcasts, documentaries, e-learning courses and corporate training that need multi-voice narration with emotional direction.

Platform

Kairos PlatformPlatform

Evaluation platform for voice agents: it builds evaluation suites from real-world use cases, simulates agent-to-agent and human-to-agent conversations against any voice model, and tracks regressions over time.

No counterpart

Cerence sells these in a stack layer with no product recorded for Hume AI yet — nothing on the other side to compare them against.

AI agent

Dealer Assist AgentAI agent

Handles dealership customer inquiries as part of Cerence's automotive AI agents portfolio.

Developer tool

Cerence Voice LabDeveloper tool

Self-service cloud portal where developers upload their own audio to benchmark Cerence's speech recognition engines on word error rate, latency and confidence, and audition and tune the text-to-speech voice library in the browser.

Hume AI sells these in a stack layer with no product recorded for Cerence yet — nothing on the other side to compare them against.

API service

Empathic Voice Interface (EVI)API service

Empathic Voice Interface (EVI) is Hume's real-time speech-to-speech voice AI API, described in Hume's docs as 'an advanced, real-time emotionally intelligent voice AI.' The product page markets it as a way to 'Build voice-first AI experiences that understand and respond to human emotions' and states it 'works seamlessly with Claude, GPT, Gemini, Grok, Kimi K2, Llama, and more' - layering voice on top of third-party language models rather than supplying its own. It is billed per conversation minute on Hume's public plans.

Expression Measurement APIAPI service

The Expression Measurement API measures emotional and vocal expression in speech, marketed as 'Measure expression in voice, offline or in real time.' Hume describes two endpoints on the one page: 'The Tagger API returns 600+ expression and voice dimensions from any audio, at scale. The Prosody API delivers real-time emotional expressions during live conversations.' Measured dimensions include emotion categories, speaking styles, vocal qualities such as pace, warmth and hesitation, and expression intensity with confidence scores.

Human Feedback APIAPI service

The Human Feedback API sells human evaluation of voice and conversational AI as an API call, marketed as 'One API call. Real human ratings.' and 'Plug human evaluation into your pipeline. Audio in, structured results out.' It supports three study types - live conversations with AI models, single-sample audio ratings with custom questions, and side-by-side A/B comparisons - with Hume handling participant recruitment, quality control and fraud detection. Hume states it powers 'evaluation for Hume's own frontier voice models.'

Octave TTSAPI service

Octave TTS is Hume's text-to-speech API, which Hume's documentation calls 'the first text-to-speech system built on LLM intelligence.' The marketing page positions it as 'Text-to-speech with emotional intelligence' that generates 'expressive, natural-sounding speech that conveys the full range of human emotion,' with SDKs for Python, TypeScript, .NET and Swift and streaming audio output. It is billed by characters on Hume's public self-serve plans.

Data service

Data SolutionsData service

Data Solutions is Hume's licensed AI training-data offering, headlined 'Voice Training Data Built by Researchers, for Researchers' and described as 'Datasets for creating realistic voices across global languages, powering our own state-of-the-art models, and now available to power yours.' The catalogue spans voice datasets covering 50+ languages, 48+ emotions and 600+ voice descriptors, plus expression and multimodal datasets for facial expression, vocal bursts, speech prosody and text emotion. Hume also sells custom data collection with defined speakers and recording conditions, and API access for programmatic data refresh and generation.