Deepgram
Developer-first voice AI API for speech-to-text, text-to-speech, and real-time voice agents
Speech-to-text API with real-time and batch transcription across 56+ languages, plus text-to-speech and voice agent endpoints
Speechmatics is a Cambridge, UK speech AI vendor that sells transcription as a metered API rather than a seat subscription. The Free plan renews 3,000 minutes (50 hours) of speech-to-text plus 1 million text-to-speech characters every month. Pro requires no commitment and bills per transcribed hour: batch Melia 1 at $0.24, batch Enhanced at $0.75, real-time Enhanced at $0.80. On-premises deployment, custom models, and audio alignment are Enterprise-only.
Speechmatics is a UK speech AI company, founded in 2006 and based in Cambridge, that sells transcription and voice synthesis as metered APIs rather than as a seat-based product. The platform covers three surfaces: real-time speech-to-text that returns output in under a second, batch transcription for pre-recorded files, and a text-to-speech endpoint aimed at conversational agents. Transcription runs across 56 or more languages and dialects, with 69 supported translation pairs, and includes speaker diarization, word-level timestamps, automatic language identification, custom dictionaries, and multi-channel handling as standard rather than as paid add-ons.
Three transcription models sit behind the same API. Enhanced is the highest-accuracy option at $0.75 per hour for batch and $0.80 per hour for real-time; Standard trades some accuracy for cost and turnaround at $0.45 per hour on both paths; and Melia 1, announced in June 2026, transcribes multilingual audio including speakers who switch language mid-sentence, at $0.24 per hour for batch. Melia 1 is currently batch-only and carries a production preview label. Optional speech-intelligence bolt-ons are billed separately per hour, from $0.12 for summaries and sentiment up to $0.65 for translation. Text-to-speech is charged at $0.011 per 1,000 characters and is English-only for now.
Speechmatics is aimed at developers and product teams embedding transcription into their own software, not at individuals looking for a meeting-notes app. The pricing shape reflects that: there is no monthly seat fee, the Free plan renews 3,000 minutes of speech-to-text every month so evaluation is genuinely open-ended, and the Pro plan simply meters hours with no commitment. The tradeoff is that everything relating to control — on-premises deployment, custom models, unlimited concurrency, audio alignment — is gated behind an Enterprise conversation.
Starting price: $0.129/hr · Free tier: yes · Model: usage-based
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Free | Free | 3,000 speech-to-text minutes (50 hours) per month, renewing · Split as 20 hours real-time and 30 hours batch · 1 million text-to-speech characters (~20 hours) per month · 2 concurrent real-time sessions, 1 batch job per second · 56+ languages, diarization, custom dictionary, language ID included · No credit card required to start |
| Pro | from $0.129/hr | Batch: Melia 1 $0.24/hr, Standard $0.45/hr, Enhanced $0.75/hr · Real-time: Standard $0.45/hr, Enhanced $0.80/hr · Text-to-speech at $0.011 per 1,000 characters · Bolt-ons per hour: translation $0.65, chapters $0.40, topics $0.20, summaries and sentiment $0.12 · Same free monthly allowance applies before metered usage · 50 concurrent real-time sessions, 10 batch jobs per second · 20% automatic discount above 500 hours per month per STT type · Billed to the second, monthly in arrears; capped at 6,000 hours per month |
| Enterprise | Custom | Volume discounts starting from 24,000 hours per year · No rate limits; unlimited concurrency · Private cloud, container, virtual appliance, and on-device deployment · Custom acoustic models, custom voice and language development · Audio alignment and early access to new features · Dedicated Customer Success Manager and Solutions Engineer |
| Pros | Cons |
|---|---|
| The free allowance renews monthly rather than being a one-off trial: 50 hours of speech-to-text plus 1 million TTS characters, no card required to begin | The advertised 'from $0.129/hr' headline is a discounted rate; the vendor's own plan-comparison table lists batch Melia 1 at $0.24/hr, and the 20% volume discount only begins above 500 hours per month |
| Pro has no seat fee and no minimum commitment — usage is billed to the second at the per-hour rate | Text-to-speech is English-only at present, so a multilingual voice agent still needs a second synthesis vendor |
| Core accuracy features are not paywalled: diarization, custom dictionary, language identification, timestamps, and audio events are available on the Free plan | Melia 1 is batch-only and labelled a production preview — real-time multilingual code-switching is not generally available yet |
| 56+ transcription languages and 69 translation pairs, with unusually deep accent and dialect coverage | On-premises, container, and on-device deployment, custom models, and audio alignment are all Enterprise-only with no self-serve path |
| Data residency choice across US, EU, and Australia regions, with ISO 27001, SOC 2 Type II, and HIPAA compliance claimed | Pro is capped at 6,000 hours per month, forcing a sales conversation for higher volumes |
| The official pricing FAQ still states 8 free hours per month while the plan cards and comparison table say 50 hours — an unresolved contradiction on the vendor's own page |
Developer-first voice AI API for speech-to-text, text-to-speech, and real-time voice agents
Voice AI API for developers: speech-to-text, speaker diarization, audio intelligence, and LLM-over-audio with per-hour usage billing
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
Low-latency voice AI platform behind the Sonic text-to-speech model, built for real-time agents
AI meeting assistant that transcribes, summarizes and extracts action items from conversations
Yes. The Free plan renews 3,000 speech-to-text minutes (50 hours) each month, split as 20 hours real-time and 30 hours batch, plus 1 million text-to-speech characters. It permits 2 concurrent real-time sessions and 1 batch job per second, and no credit card is needed to begin. Note that the vendor's own pricing FAQ still quotes 8 free hours per month, a contradiction they have not resolved.
Speechmatics bills per hour of audio on the Pro plan. Its published plan-comparison table lists batch Melia 1 at $0.24, batch Standard at $0.45, batch Enhanced at $0.75, real-time Standard at $0.45, and real-time Enhanced at $0.80 per hour. Text-to-speech is priced at $0.011 per 1,000 characters. The $0.129 per hour headline on the pricing page is a discounted Melia rate rather than the list price.
Melia 1 is the multilingual transcription model Speechmatics announced in June 2026. It handles speakers who switch language mid-conversation without requiring you to select a language first, and covers the full set of 56+ languages while matching the Standard model on accuracy. It is currently batch-only and carries a production preview label, with the company stating real-time support will arrive in preview first.
Only on the Enterprise plan. Enterprise unlocks private cloud, container, virtual appliance, and on-device deployment along with GPU and CPU based models. The Free and Pro plans are SaaS-only, though both let you pick a US, EU, or Australia cloud region. Custom acoustic models and custom voice development are also Enterprise-only and require contacting sales for a quote.
Yes. A 20% discount applies automatically to billable usage above 500 hours per month for each speech-to-text type, so real-time enhanced and real-time standard hours are counted separately rather than pooled. Further discounts begin from 24,000 hours of usage per year and are negotiated directly. Pro usage is capped at 6,000 hours per month, so higher volumes require an Enterprise contract.
Speechmatics transcribes 56 or more languages and dialects, including Arabic, Mandarin, Hindi, Welsh, Swahili, and Uyghur, and supports 69 translation pairs across roughly 30 target languages. Text-to-speech is much narrower: it is English-only today, with the company saying more languages are coming. Teams building multilingual voice agents will therefore need a separate speech synthesis provider for now.
Pro customers are billed on the first of each month for the previous month's usage, charged to the second based on the per-hour rate, with no seat fee or minimum commitment. Adding a card is only required once you exceed the free monthly allowance. Enterprise billing is arranged case by case, and early-stage companies can apply to a startup program offering credits.
It targets that use case directly. Real-time transcription returns output in under a second, and the platform bundles text-to-speech and speaker recognition alongside it. Concurrency is the practical constraint: the Free plan allows 3 concurrent voice agent conversations and Pro allows 6, with unlimited concurrency reserved for Enterprise. Production deployments serving many simultaneous callers will need an Enterprise agreement.