ElevenLabs
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
Empathic voice AI for developers — Octave 2 expressive text-to-speech and the EVI speech-to-speech API with per-character and per-minute pricing
Hume AI is a developer platform for expressive voice AI: the Octave text-to-speech model (Octave 2 in preview) and the EVI speech-to-speech API. The free tier includes 10,000 TTS characters and 5 EVI minutes per month; paid plans run from $3 to $500/month, with overage falling from $0.15 to $0.05 per 1,000 TTS characters and $0.07 to $0.04 per EVI minute as tiers rise. Best for developers building voice agents, narration, and character voices.
Hume AI is a voice-AI research lab and developer platform built around emotional expression. Its two commercial products are Octave, a text-to-speech system built on an LLM that interprets what the text means before deciding how to say it — with voice design from plain-text prompts, voice cloning, and voice conversion — and EVI, the Empathic Voice Interface, a speech-to-speech API that listens to a user’s vocal modulation and replies expressively with native interruption handling and back-channeling. As of July 2026 the current generations are Octave 2 and EVI 4-mini, both in preview alongside the generally available Octave 1 and EVI 3. The company also publishes TADA, an open-source LLM TTS system, and its research base covers 50+ languages, 48+ emotions, and 600+ voice descriptors.
Billing is usage-metered inside monthly subscriptions: every tier includes a TTS character allowance and an EVI minute allowance, with overage rates that drop as tiers rise — from $0.15 down to $0.05 per 1,000 TTS characters and from $0.07 down to $0.04 per EVI minute. Plans run from a permanent free tier (10,000 characters and 5 EVI minutes per month) through Starter at $3, Creator at $14, Pro at $70, Scale at $200, and Business at $500 per month, with custom Enterprise deals adding SOC 2 Type II, GDPR, and HIPAA compliance. Developers integrate through REST and WebSocket APIs plus SDKs for React, TypeScript, Python, iOS/macOS, and .NET; voice cloning is unlimited on every tier.
Hume AI is aimed at people who build with APIs rather than consumers looking for a point-and-click voice app — the platform playground exists for testing, but the product is the API surface. The free tier and $3 Starter plan make evaluation nearly free, while the per-minute EVI pricing model maps cleanly onto conversational products that bill their own users by usage.
Starting price: $3/mo · Free tier: yes · Model: freemium
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Free | $0 | 10,000 TTS characters/month (~10 minutes) · 5 EVI minutes/month · 1 concurrent connection · Unlimited voice cloning (create and use) · Discord support |
| Starter | $3/mo | 30,000 TTS characters/month (~30 minutes) · 40 EVI minutes/month, then $0.07/minute · 5 concurrent connections · 50% off the first month |
| Creator | $14/mo | 140,000 TTS characters/month (~140 minutes), then $0.15/1,000 characters · 200 EVI minutes/month, then $0.07/minute · 5 concurrent connections |
| Pro | $70/mo | 1,000,000 TTS characters/month (~1,000 minutes), then $0.12/1,000 characters · 1,200 EVI minutes/month, then $0.06/minute · 10 concurrent connections |
| Scale | $200/mo | 3,300,000 TTS characters/month (~3,300 minutes), then $0.10/1,000 characters · 5,000 EVI minutes/month, then $0.05/minute · 20 concurrent connections · 3 team seats |
| Business | $500/mo | 10,000,000 TTS characters/month (~10,000 minutes), then $0.05/1,000 characters · 12,500 EVI minutes/month, then $0.04/minute · 30 concurrent connections · 5 team seats |
| Enterprise | Custom | Custom volumes and unlimited team seats · SOC 2 Type II, GDPR, HIPAA compliance · Slack support channel |
| Pros | Cons |
|---|---|
| Cheap to prototype: a permanent free tier plus a $3/month Starter plan with 40 EVI minutes included | Developer-first product — there is no polished consumer app; you build against APIs and SDKs or work in the platform playground |
| Business-tier TTS overage of $0.05/1,000 characters undercuts most premium expressive-TTS rivals at scale | The newest models (Octave 2, EVI 4-mini) are labeled preview, so behavior and rates may still shift |
| EVI handles interruptions and back-channeling natively — conversational behaviors that must be hand-built on plain TTS stacks | Concurrency caps (1 connection on Free, 30 on Business) constrain production voice apps without an Enterprise deal |
| Voice cloning is unlimited on every tier, including Free | High-volume conversational usage adds up: 10,000 EVI minutes beyond the Pro allowance costs about $600 at $0.06/minute |
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
Low-latency voice AI platform behind the Sonic text-to-speech model, built for real-time agents
AI voice platform for voice cloning, text-to-speech, and real-time voice agents, plus deepfake detection and audio watermarking
Developer platform to build, test, and deploy AI voice agents for phone calls with usage-based per-minute pricing
AI text-to-speech studio with 200+ voices across 35+ languages for voiceovers, dubbing, and voice agents
There is a permanent free tier with 10,000 text-to-speech characters (~10 minutes) and 5 EVI conversation minutes per month, one concurrent connection, and unlimited voice cloning. It is enough to evaluate the APIs but not to run anything in production.
Paid plans are Starter $3/month, Creator $14/month, Pro $70/month, Scale $200/month, and Business $500/month, each with monthly TTS-character and EVI-minute allowances. Overage falls from $0.15 to $0.05 per 1,000 TTS characters and $0.07 to $0.04 per EVI minute as tiers rise. Enterprise pricing is custom.
EVI (Empathic Voice Interface) is Hume's speech-to-speech API: it listens, interprets vocal expression, and replies in an expressive voice with native interruption handling and back-channeling. EVI 3 is generally available and EVI 4-mini is in preview with expanded language support and lower latency.
Octave is Hume's text-to-speech system built on an LLM, so it interprets the meaning of the text rather than just pronouncing it. It supports voice design from text prompts, voice cloning, and voice conversion. Octave 2, the current generation, is available in preview.
Yes. Voice cloning — both creating and using cloned voices — is unlimited on every plan, including the free tier. Commercial usage rights are tied to paid plans per the pricing page.
Both offer expressive TTS and conversational voice APIs. Hume's differentiators are its emotion-science research base (48+ emotions, 600+ voice descriptors), the speech-to-speech EVI API with built-in interruption handling, and aggressive at-scale pricing — $0.05 per 1,000 characters overage on the $500/month Business plan.
Official SDKs cover React, TypeScript/Node.js, Python (sync and async clients), iOS/macOS, and .NET, alongside REST and WebSocket APIs documented on the developer portal.