AssemblyAI
Voice AI API for developers: speech-to-text, speaker diarization, audio intelligence, and LLM-over-audio with per-hour usage billing
Developer-first voice AI API for speech-to-text, text-to-speech, and real-time voice agents
Deepgram is a developer-first voice AI platform for speech-to-text, text-to-speech, and voice agents, billed by usage with no minimums. New accounts get $200 in free credits. Nova-3 speech-to-text runs about $0.0077 per minute for pre-recorded audio, Aura-2 text-to-speech about $0.030 per 1,000 characters, and the Voice Agent API starts near $0.075 per minute. Best for engineering teams building real-time or batch voice features via API rather than through a consumer app.
Deepgram is a voice AI platform aimed at developers and product teams rather than end users. It exposes three main capabilities through one API: speech-to-text (the Nova-3 family, plus the voice-agent-focused Flux model), text-to-speech (the Aura-2 and Aura-1 voices), and a unified Voice Agent API that stitches transcription, a language model, and speech synthesis into a single real-time stream. Everything is billed by usage — per minute for transcription, per 1,000 characters for synthesis, and per minute for voice agents — with no minimums and $200 in free credits to start.
Deepgram’s positioning centers on latency and deployment flexibility. Streaming transcription lands in roughly the 200-300ms range, which is fast enough for live captioning and conversational agents, and Nova-3 covers monolingual English plus a multilingual model spanning more than 30 languages. Beyond the cloud API, Deepgram offers real-time streaming and self-hosted, on-premises deployment for teams with data-residency or compliance constraints — a differentiator against API-only competitors. Optional add-ons handle PII redaction, speaker diarization, and keyterm prompting for domain vocabulary.
Deepgram is built for engineers embedding voice into a product, not for individuals who want a finished app. The $200 free credit is enough to prototype a transcription pipeline or a voice agent, but the real value shows up at production volume where per-minute costs of fractions of a cent compound into meaningful savings versus higher-priced APIs.
Starting price: $0 · Free tier: yes · Model: usage-based
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Pay As You Go | Usage-based | $200 in free credits on signup, no credit card, no expiration · No minimums or commitments — pay only for audio processed · Nova-3 speech-to-text from ~$0.0048/min streaming, ~$0.0077/min pre-recorded · Aura-2 text-to-speech ~$0.030 per 1,000 characters |
| Growth | Prepaid credits | Prepaid annual credits save roughly 20% versus Pay As You Go · Lower per-minute STT and per-character TTS rates · Same models and API as Pay As You Go |
| Enterprise | Custom | Custom volume pricing and dedicated capacity · Self-hosted / on-premises deployment · SOC 2 and compliance controls, custom model training · Priority support and SLAs |
| Pros | Cons |
|---|---|
| Usage-based pricing with no minimums and $200 in free credits to start | Developer-only: there is no consumer app, so every use requires engineering integration |
| Low streaming latency (roughly 200-300ms) suits real-time voice agents | Aura text-to-speech voices are functional but less expressive than ElevenLabs for premium work |
| Self-hosting, SOC 2, and on-prem options make it viable for regulated, high-volume workloads | Multilingual and pre-recorded rates cost noticeably more than the headline monolingual streaming price |
| One API spans speech-to-text, text-to-speech, and full voice agents |
Voice AI API for developers: speech-to-text, speaker diarization, audio intelligence, and LLM-over-audio with per-hour usage billing
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
AI voice platform for voice cloning, text-to-speech, and real-time voice agents, plus deepfake detection and audio watermarking
AI text-to-speech studio with 200+ voices across 35+ languages for voiceovers, dubbing, and voice agents
Deepgram is usage-based rather than free, but new accounts receive $200 in free credits with no credit card and no expiration. The Pay As You Go plan has no minimums, so you only pay for the audio you process once the credit runs out. Speech-to-text starts around $0.0077 per minute.
Nova-3, Deepgram's flagship speech-to-text model, costs roughly $0.0048 per minute for streaming and $0.0077 per minute for pre-recorded audio on Pay As You Go. The Growth tier lowers these rates by about 20% through prepaid annual credits, and multilingual transcription costs slightly more.
Both are speech-to-text APIs, but Deepgram competes on low latency and self-hosting for real-time voice agents, while AssemblyAI leans toward higher-level audio intelligence and LLM features. Deepgram also ships its own Aura text-to-speech and a unified Voice Agent API in the same platform.
Yes. Deepgram's Aura-2 text-to-speech model costs about $0.030 per 1,000 characters, and the older Aura-1 costs about $0.015. The voices are tuned for real-time voice agents and priced well below premium voice providers, though they are generally less expressive than ElevenLabs.
Yes. Deepgram supports cloud, real-time streaming, and self-hosted on-premises deployment. Self-hosting is available on the Enterprise plan with custom pricing and appeals to teams with strict data-residency, compliance, or high-throughput requirements that a shared cloud API cannot satisfy.
Deepgram's Nova-3 multilingual model transcribes more than 30 languages, including Spanish, French, German, Japanese, Korean, Hindi, and Mandarin. Multilingual transcription is priced slightly higher than the English-only monolingual model, at roughly $0.0058 per minute for streaming.