Deepgram
Developer-first voice AI API for speech-to-text, text-to-speech, and real-time voice agents
Speech-to-text API with real-time and async transcription across 100+ languages, with audio intelligence bundled into the per-hour rate
Gladia is a Paris-based speech-to-text API built on its Solaria models, sold by the hour of audio rather than by seat. The Starter plan is pay-as-you-go at $0.61 per hour for async transcription and $0.75 per hour for real-time, with 10 hours free every month and no card required. Growth drops to as low as $0.20 per hour async but needs an upfront commitment and a sales conversation. The per-hour rate sits above several rivals, but diarization, language detection, and 100+ languages are included rather than billed as add-ons.
Gladia is a French speech-to-text company, founded in Paris in 2022, that sells transcription as a metered API rather than as a finished app. Two endpoints cover the workload: async transcription for recordings and long-form audio, and real-time streaming that returns output in under 300 milliseconds, with partial transcripts arriving in under 100 milliseconds. Behind both sits the Solaria model family. Solaria-1 is the universal option, covering more than 100 languages and handling speakers who switch language mid-sentence. Solaria-3, released in June 2026, trades that breadth for accuracy on real production audio in English and core European languages, and is aimed at noisy, fast-paced business recordings such as customer calls.
The pricing shape is the interesting part. Where several competitors meter each intelligence feature separately, Gladia folds automatic language detection, speaker diarization, word-level timestamps, and the full language set into a single per-hour rate. The self-serve Starter plan bills $0.61 per hour for async and $0.75 per hour for real-time, with 10 hours free every month, 30 concurrent real-time requests, and 25 concurrent async requests. Growth advertises rates as low as $0.20 per hour but gates them behind an upfront commitment and a sales conversation. Beyond raw transcription, the platform adds PII redaction, sentiment analysis, summarization, and an Audio to LLM path that runs custom prompts over a transcript in the same API call. Billing is moving to a prepaid credit wallet as of July 2026, with per-hour rates stated as unchanged.
Gladia is built for engineering teams embedding transcription into their own product, not for individuals who want a meeting-notes app. The headline rate is higher than several rivals, so the honest test is whether you need the bundled audio intelligence: if you only want raw English transcripts at volume, cheaper per-hour options exist. If you need diarized, multilingual output with sentiment or entity extraction attached, the bundled rate closes most of that gap. One caveat worth weighing before committing: OVH Groupe entered exclusive negotiations to acquire Gladia in June 2026, and that deal had not been announced as closed at the time of writing.
Starting price: $0 · Free tier: yes · Model: usage-based
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Starter | $0.61/hr async, $0.75/hr real-time | 10 hours of transcription free every month, renewing · Async transcription $0.61/hr, real-time streaming $0.75/hr · Pay-as-you-go with immediate self-serve signup, no sales call · 30 concurrent real-time requests, 25 concurrent async requests · Automatic language detection and switching, speaker diarization, 100+ languages · GDPR, HIPAA, and AICPA SOC 2 Type 2 coverage · Support via help center and Discord |
| Growth | from $0.20/hr async, $0.25/hr real-time | Roughly 67% below Starter rates, per the vendor's own pricing page · Requires an upfront volume commitment negotiated with sales · Flexible concurrent request limits and custom volume discounts · Automatic opt-out from model training · 99.9% uptime SLA and priority processing queue |
| Enterprise | Custom | Custom models, fine-tuning, and debundled per-feature pricing · Unlimited concurrent requests and dedicated infrastructure · Zero data retention and model training opt-out by default · Custom hosting options and 99.9% uptime SLA · Dedicated Slack channel, account manager, and premium support |
| Pros | Cons |
|---|---|
| Core capabilities are bundled into the per-hour rate rather than sold as separate line items — diarization, language detection, and all 100+ languages are included on the entry plan | Starter async at $0.61 per hour is meaningfully above rivals' entry rates — Deepgram's pre-recorded Nova-3 works out near $0.46 per hour and Speechmatics lists batch Melia 1 at $0.24 per hour — so the bundling only pays off if you actually use diarization, sentiment, or entity extraction |
| The 10-hour monthly free allowance renews rather than expiring as a one-off trial, and needs no credit card | The headline $0.20 per hour Growth rate is not self-serve: it requires an upfront volume commitment and a sales conversation, so the published number is unreachable for small teams |
| Starter is genuinely self-serve at 30 concurrent real-time and 25 concurrent async requests, which is enough concurrency to run a small production workload without talking to sales | OVH Groupe entered exclusive negotiations to acquire Gladia on 11 June 2026 and the transaction had not been announced as closed as of 20 July 2026, leaving roadmap and pricing continuity unresolved |
| Two distinct models let you pick breadth (Solaria-1, 100+ languages) or accuracy on business audio (Solaria-3, English and core European languages) | Gladia sells speech-to-text only — there is no text-to-speech, so a full voice agent still needs a second synthesis vendor |
| Audio to LLM removes the glue code between transcription and a language model for summaries and extraction | The uptime SLA and priority processing queue start at Growth; Starter customers get no contractual availability guarantee |
| Custom hosting, fine-tuning, and zero data retention are Enterprise-only with no published price |
Developer-first voice AI API for speech-to-text, text-to-speech, and real-time voice agents
Speech-to-text API with real-time and batch transcription across 56+ languages, plus text-to-speech and voice agent endpoints
Voice AI API for developers: speech-to-text, speaker diarization, audio intelligence, and LLM-over-audio with per-hour usage billing
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
Low-latency voice AI platform behind the Sonic text-to-speech model, built for real-time agents
Gladia has a free allowance rather than a free product. Signing up gives you up to 10 hours of transcription each month at no charge, renewing monthly, with no credit card required. Beyond that allowance the Starter plan meters usage at $0.61 per hour for async transcription and $0.75 per hour for real-time streaming.
On the self-serve Starter plan, async transcription costs $0.61 per hour of audio and real-time streaming costs $0.75 per hour. The Growth plan advertises rates as low as $0.20 per hour async and $0.25 per hour real-time, roughly 67% below Starter, but requires an upfront commitment negotiated with sales. Enterprise pricing is custom.
Solaria-1 is the universal model, covering more than 100 languages and handling speakers who switch language mid-conversation. Solaria-3, released in June 2026, is narrower but more accurate on real production audio in English and core European languages such as French, Italian, Spanish, and German. Pick Solaria-1 for breadth and Solaria-3 for business call accuracy.
OVH Groupe announced on 11 June 2026 that it had entered exclusive negotiations to acquire Gladia, citing sovereign voice AI ambitions for OVHcloud and OVHai. As of 20 July 2026 no completion had been announced. Gladia continues to sell its own plans, publish its own pricing, and ship product updates independently.
No. Gladia is a speech-to-text and audio intelligence provider only. It covers async and real-time transcription plus features like diarization, sentiment analysis, PII redaction, and summarization, but it does not synthesize speech. Teams building a full voice agent will need to pair Gladia with a separate text-to-speech vendor.
Gladia announced in July 2026 that it is moving from end-of-month invoicing to a prepaid credit wallet. You buy credits upfront, watch the balance in the developer console, and top up manually or automatically when it drops below a threshold. Per-hour rates were stated as unchanged by the switch, and payments run through Stripe.
More than 100 languages are supported on every paid plan, including automatic language detection and switching within a single recording. Language coverage is not tiered, so an entry-level Starter account gets the same language set as Enterprise. Solaria-3 narrows to English and core European languages in exchange for higher accuracy on noisy business audio.
Not on the headline rate. Gladia's self-serve async rate is $0.61 per hour, while Deepgram's pre-recorded Nova-3 model works out to roughly $0.46 per hour. Gladia becomes competitive when you need diarization, sentiment, or entity extraction, since those are bundled into its per-hour price rather than billed as separate add-ons.