Descript
Text-based video and podcast editor with AI transcription, voice cloning, and filler-word removal
Voice AI API for developers: speech-to-text, speaker diarization, audio intelligence, and LLM-over-audio with per-hour usage billing
AssemblyAI is a developer API for speech-to-text and voice AI — async transcription, real-time streaming, speaker diarization, and LLM-over-audio. Pre-recorded transcription starts at $0.15 per hour and streaming at $0.15 per hour, and new accounts get $50 in free credits with no card. Its Universal-2 model covers 99 languages. Best for engineers adding transcription or voice agents to products; it is API-only with no consumer app.
AssemblyAI is a developer-focused API for turning speech into text and building voice-driven features. Its core is high-accuracy transcription — both asynchronous, for pre-recorded files, and real-time, for live streaming audio — delivered through the Universal model family. On top of raw transcription it layers speaker diarization, audio intelligence (entity and topic detection, sentiment, summarization, content moderation), and PII redaction, so a single integration can produce structured, searchable output rather than a plain transcript.
More recently the product has expanded into voice AI infrastructure. An LLM Gateway lets you run large language models like Claude and Gemini directly over transcribed audio for question-answering and summarization, and a Voice Agent API bundles transcription, reasoning, speech, and turn detection for conversational applications. Pricing is usage-based: pre-recorded transcription starts at $0.15 per hour, streaming at $0.15 per hour, with add-ons billed on top. New accounts receive $50 in free credits with no credit card, and enterprise plans add HIPAA, SOC 2 Type 2, ISO 27001, and GDPR compliance at no premium.
AssemblyAI is for engineers and product teams building audio features into their own software. It is deliberately API-only, so it rewards teams that can write code and want control, not end users looking for a finished app.
Starting price: $0.15/hr · Free tier: yes · Model: usage-based
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Free | $50 in credits | $50 in free credits, no credit card required · Covers roughly 185 hours of pre-recorded audio · Limited to 5 new streams per minute |
| Pay-as-you-go | from $0.15/hr | Universal-2 async transcription $0.15/hr; Universal-3 Pro $0.21/hr · Streaming from $0.15/hr; up to 100 new streams per minute · Speaker diarization add-on +$0.02/hr async, +$0.12/hr streaming · Access to all models and audio-intelligence add-ons |
| Enterprise | Custom | HIPAA BAA, SOC 2 Type 2, ISO 27001, GDPR at no premium · Higher concurrency via contract · Volume pricing and dedicated support |
| Pros | Cons |
|---|---|
| Transparent per-hour usage pricing with no upfront commitment | API and SDK only — there is no web app or no-code UI, so non-developers cannot use it directly |
| Generous $50 free signup credit with no credit card required | Audio-intelligence features are add-ons that stack on the base rate and raise effective cost |
| Broad language coverage (99 languages on Universal-2) plus 100+ translation targets | The newer Universal-3 Pro currently supports far fewer languages than Universal-2's 99 |
| One API spans async, streaming, audio intelligence, voice agents, and LLM-over-audio | In-region model pricing rises 10% from July 1, 2026 unless requests use the global region |
| Enterprise compliance (HIPAA BAA, SOC 2 Type 2, ISO 27001, GDPR) included at no premium | Enterprise pricing is contact-sales only, not published |
Text-based video and podcast editor with AI transcription, voice cloning, and filler-word removal
AI meeting assistant that transcribes, summarizes and extracts action items from conversations
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
AI meeting assistant that transcribes, summarizes, and searches your conversations
Pre-recorded transcription starts at $0.15/hr on Universal-2 and $0.21/hr on Universal-3 Pro. Streaming starts at $0.15/hr, billed by audio duration with audio-intelligence add-ons priced separately.
Yes. New accounts get $50 in free credits with no credit card required, which covers roughly 185 hours of pre-recorded transcription, capped at 5 new streams per minute.
The Universal-2 model supports 99 languages for pre-recorded audio, and translation supports 100+ target languages. The newer Universal-3 Pro currently covers a smaller language set.
Yes. Speaker diarization is a stackable add-on, priced at an extra $0.02/hr for pre-recorded audio and $0.12/hr for streaming audio.
AssemblyAI is an API-first product with official Python and JavaScript SDKs. There is no consumer web or mobile app — it is meant to be integrated into your own software.
Yes. Its LLM Gateway lets you run models such as Claude and Gemini over transcribed audio, billed per million input and output tokens.