Octave TTS
Octave TTS is Hume's text-to-speech API, which Hume's documentation calls 'the first text-to-speech system built on LLM intelligence.' The marketing page positions it as 'Text-to-speech with emotional intelligence' that generates 'expressive, natural-sounding speech that conveys the full range of human emotion,' with SDKs for Python, TypeScript, .NET and Swift and streaming audio output. It is billed by characters on Hume's public self-serve plans.
Find alternatives to Octave TTS
Bought on its own on Hume's public plan ladder: Free $0/month, Starter $3/month, Creator $7/month, Pro $70/month, Scale $200/month, Business $500/month, Enterprise 'Custom'. Text-to-speech is metered in characters on top of the plan allowance: Creator includes '140,000(~140 minutes)' monthly characters with overage at '$0.15/1,000'; Pro includes '1,000,000(~1,000 minutes)' with overage at '$0.12/1,000'. The '$0.15/1,000' Creator overage was independently re-confirmed by the verification pass 2026-09-05.
What it does
- Text to speech
How it compares
-
Voice Design
For each voice description and text input, we then generated three samples using Octave and three using Elevenlabs Voice Design. Raters (N = 180) were instructed to blindly compare paired Octave and Elevenlabs speech samples generated using the same prompt.
Official documentation · 20 Sep 2026 -
Sonic / Ink
Cartesia | Hume AI alternative for voice agents
Company disclosure · 20 Sep 2026
Sources
- Pricing
-
Text-to-Speech row lists 'Octave 1' and 'Octave 2'.
- Description
-
Official documentation · 5 Sep 2026
Page title 'Octave - Text-to-Speech with Emotional Intelligence | Hume AI'; 'Text-to-speech with emotional intelligence. Generate expressive, natural-sounding speech that conveys the full range of human emotion.'
- Sold within
-
hume.ai/pricing (already cited pricing-page elsewhere on this record); a first-class line item on Hume's own pricing page, bought directly on the plan ladder.
- Name
-
Official documentation · 5 Sep 2026
The docs draw the product/model line themselves: 'Octave TTS is the first text-to-speech system built on LLM intelligence' (the product), while 'Octave' is separately described as 'a speech-language model'. The docs spelling 'Octave TTS' was preferred because it disambiguates the product from the model family.
Also from Hume AI
-
Kairos Platform Platform
Evaluation platform for voice agents: it builds evaluation suites from real-world use cases, simulates agent-to-agent and human-to-agent conversations against any voice model, and tracks regressions over time.
-
Empathic Voice Interface (EVI) API service
Empathic Voice Interface (EVI) is Hume's real-time speech-to-speech voice AI API, described in Hume's docs as 'an advanced, real-time emotionally intelligent voice AI.' The product page markets it as a way to 'Build voice-first AI experiences that understand and respond to human emotions' and states it 'works seamlessly with Claude, GPT, Gemini, Grok, Kimi K2, Llama, and more' - layering voice on top of third-party language models rather than supplying its own. It is billed per conversation minute on Hume's public plans.
-
Creator Studio Application
Creator Studio is Hume's web application, hosted at app.hume.ai, for turning written documents into narrated audio: 'From document to audio, effortlessly. Import your content, assign voices, add acting directions, and export production-ready audio.' Hume targets it at creators producing audiobooks, videos, podcasts, documentaries, e-learning courses and corporate training that need multi-voice narration with emotional direction.
-
Expression Measurement API API service
The Expression Measurement API measures emotional and vocal expression in speech, marketed as 'Measure expression in voice, offline or in real time.' Hume describes two endpoints on the one page: 'The Tagger API returns 600+ expression and voice dimensions from any audio, at scale. The Prosody API delivers real-time emotional expressions during live conversations.' Measured dimensions include emotion categories, speaking styles, vocal qualities such as pace, warmth and hesitation, and expression intensity with confidence scores.
-
Human Feedback API API service
The Human Feedback API sells human evaluation of voice and conversational AI as an API call, marketed as 'One API call. Real human ratings.' and 'Plug human evaluation into your pipeline. Audio in, structured results out.' It supports three study types - live conversations with AI models, single-sample audio ratings with custom questions, and side-by-side A/B comparisons - with Hume handling participant recruitment, quality control and fraud detection. Hume states it powers 'evaluation for Hume's own frontier voice models.'
-
Data Solutions Data service
Data Solutions is Hume's licensed AI training-data offering, headlined 'Voice Training Data Built by Researchers, for Researchers' and described as 'Datasets for creating realistic voices across global languages, powering our own state-of-the-art models, and now available to power yours.' The catalogue spans voice datasets covering 50+ languages, 48+ emotions and 600+ voice descriptors, plus expression and multimodal datasets for facial expression, vocal bursts, speech prosody and text emotion. Hume also sells custom data collection with defined speakers and recording conditions, and API access for programmatic data refresh and generation.
Something wrong here? Send a correction — quote this product id: hume-ai-octave-tts.