AssemblyAI
Voice AI API for developers: speech-to-text, speaker diarization, audio intelligence, and LLM-over-audio with per-hour usage billing
6 tools ranked · last updated Jul 20, 2026 · how we picked
The best speech-to-text API in 2026 is AssemblyAI at $0.15 per hour of audio, with 99 languages and HIPAA, SOC 2 and GDPR included at no premium. Speechmatics is the best free pick: 3,000 minutes — 50 hours — of transcription renewing every month with no credit card. Deepgram is the pick for real-time voice agents at roughly $0.0048 per minute streaming, and it is the only vendor here that will let you self-host. The top four are metered purely by audio processed, with no seat fee and no minimum commitment.
Prices last verified Jul 20, 2026 against official pricing pages.
AssemblyAI is the default choice for most teams adding transcription to a product, because the entry rate is the lowest here for a mature API: $0.15 per hour of audio for pre-recorded Universal-2 transcription and $0.15 per hour for streaming, with the newer Universal-3 Pro at $0.21 per hour. Universal-2 covers 99 languages and translation reaches 100+ targets, and HIPAA BAA, SOC 2 Type 2, ISO 27001 and GDPR are included at no premium rather than sold as an enterprise upsell, which is unusual at this price. New accounts get $50 in free credits with no credit card — roughly 185 hours of pre-recorded audio, capped at 5 new streams per minute. On top of raw transcription sit speaker diarization, audio intelligence, PII redaction and an LLM Gateway that runs models like Claude and Gemini over a transcript. Two caveats: add-ons stack on the base rate, with diarization at +$0.02/hr async and +$0.12/hr streaming, and Universal-3 Pro currently supports far fewer languages than Universal-2’s 99.
Deepgram is the pick when latency is the product. Nova-3 streaming transcription runs about $0.0048 per minute and lands in roughly the 200–300ms range, which is fast enough for live captioning and conversational agents, and the same platform covers Aura-2 text-to-speech at about $0.030 per 1,000 characters plus a unified Voice Agent API starting near $0.075 per minute — so one vendor and one integration can carry an entire voice pipeline. New accounts get $200 in free credits with no credit card and no expiration, the Pay As You Go plan has no minimums, and the Growth tier’s prepaid annual credits cut rates by roughly 20%. It is also the only pick here offering self-hosted, on-premises deployment, on the Enterprise plan. Two caveats: pre-recorded audio costs $0.0077 per minute, about $0.46 an hour and meaningfully above AssemblyAI’s batch rate, and multilingual transcription is priced higher again at around $0.0058 per minute streaming.
Speechmatics has the most generous free allowance in this category, and it renews rather than expiring: 3,000 speech-to-text minutes — 50 hours — every month, split as 20 hours real-time and 30 hours batch, plus 1 million text-to-speech characters, with no credit card to begin. Diarization, custom dictionary, language identification and word-level timestamps are all included on the free plan rather than paywalled. Paid usage on Pro carries no seat fee and no commitment: batch Melia 1 at $0.24/hr, batch Standard $0.45/hr, batch Enhanced $0.75/hr, real-time Standard $0.45/hr and real-time Enhanced $0.80/hr, billed to the second with a 20% automatic discount above 500 hours a month. Coverage runs to 56+ languages and 69 translation pairs, with US, EU or Australia data residency. Two caveats: the “from $0.129/hr” headline is a discounted rate, not the list price, and Melia 1 is batch-only under a production preview label.
Gladia sells the opposite pricing shape to most rivals: instead of metering each intelligence feature, it folds speaker diarization, automatic language detection, word-level timestamps and all 100+ languages into one per-hour rate. Self-serve Starter is $0.61 per hour for async transcription and $0.75 per hour for real-time, with 10 hours free every month on a renewing basis and no card required, plus 30 concurrent real-time and 25 concurrent async requests — enough to run a small production workload without a sales call. Streaming returns output under 300ms, partials under 100ms, and the Audio to LLM path runs custom prompts over a transcript in the same API call. Two caveats: $0.61/hr is well above rivals’ entry rates, so the bundle only pays off if you genuinely use diarization or sentiment; and OVH Groupe entered exclusive negotiations to acquire Gladia in June 2026, a deal not announced as closed, leaving roadmap continuity unresolved.
Groq is the outlier here: not a speech company, but an inference cloud running open models on custom LPU silicon, including Whisper Large v3 Turbo for transcription and Orpheus for synthesis. It earns a place because the free tier is genuinely free — all models, an API key from the console, no credit card — and because the LPU hardware is built for exactly the latency-sensitive voice workloads where a generic GPU backend struggles. The API is OpenAI-compatible, so migrating existing transcription code is close to a two-line change of base URL and key, and the paid Developer tier adds up to roughly 10x higher rate limits, a Batch API at 50% off async jobs and prompt caching at 50% off cached input. Two caveats: free-tier rate limits are tight for production at around 30 requests per minute, and Groq’s published headline rates are per-token figures for text models — you are buying hosted Whisper, not a purpose-built speech platform with diarization and PII redaction attached.
ElevenLabs is the pick when transcription is one half of a voice product and synthesis is the other. Its Scribe speech-to-text API sits alongside text-to-speech in 70+ languages, voice cloning, dubbing and deployable conversational agents, all behind one REST surface, which removes a second vendor from the stack. The free tier gives 10,000 credits a month with no commercial use; Starter is $6/month for 30k credits and a commercial license, Creator $22/month for 121k credits, Pro $99/month for 600k credits and 44.1kHz PCM audio via API, and Scale $299/month for 1.8M credits. Business at $990/month adds low-latency TTS from $0.05/min. Two caveats: billing is credit-based rather than per hour of audio, so transcription costs are harder to forecast than a flat $0.15/hr rate; and credit consumption on long-form, high-quality audio escalates quickly on the lower tiers.
We ranked these on the cost of an hour of audio, what arrives with that hour versus what is billed as an add-on, how usable the free allowance is for real evaluation rather than a demo, and whether latency and concurrency limits survive contact with production. Every pick is a genuine developer API with no seat fee — consumer meeting-notes apps were excluded on purpose, because a transcription endpoint you can call at 3am is a different product from a dashboard someone has to log into. Where a vendor’s headline number is a discounted or commitment-gated rate rather than what a self-serve account actually pays, we said so, since that gap is the most common way these pricing pages mislead. Pricing was verified on July 20 2026 against each product’s official pricing page; no tool paid or provided incentives to appear in this list.
Voice AI API for developers: speech-to-text, speaker diarization, audio intelligence, and LLM-over-audio with per-hour usage billing
Developer-first voice AI API for speech-to-text, text-to-speech, and real-time voice agents
Speech-to-text API with real-time and batch transcription across 56+ languages, plus text-to-speech and voice agent endpoints
Speech-to-text API with real-time and async transcription across 100+ languages, with audio intelligence bundled into the per-hour rate
Ultra-fast, low-cost inference for open models on custom LPU chips (groq.com — not xAI's Grok)
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
Every tool in this list has a full profile in our directory with pricing verified against its official pricing page on the date shown on its stamp. Ranking reflects verified pricing, free-tier generosity, platform coverage, and documented capabilities — not sponsorships. Nobody can pay to appear here. Read the full methodology.
Yes — 6 of the 6 tools here have a free tier: AssemblyAI, Deepgram, Speechmatics, Gladia, Groq, ElevenLabs. Pricing verified Jul 20, 2026.
ElevenLabs has the lowest verified monthly starting price in this list at $6/mo, checked against its official pricing page on Jul 3, 2026.
6 of the 6 tools list an API: AssemblyAI, Deepgram, Speechmatics, Gladia, Groq, ElevenLabs.