ElevenLabs
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
Voice cloning and text-to-speech platform with a shared voice marketplace and pay-as-you-go API, built on open-source Fish Speech models
Fish Audio is a voice-cloning and text-to-speech platform built on the open-source Fish Speech and OpenAudio models. The free tier gives 8,000 credits per month (about 7 minutes); paid plans start at $11/month, and the TTS API is pay-as-you-go at $15 per million UTF-8 bytes, roughly 12 hours of speech. Best for developers and creators who want low-cost, multilingual voice synthesis with instant cloning.
Fish Audio is an AI voice platform for text-to-speech and voice cloning, built on top of the open-source Fish Speech and Fish Diffusion models released through its OpenAudio research effort. Its hosted S2.1 Pro engine turns text into expressive speech in 30+ languages, supports emotion and effect tags like [angry], [whispering], and [laughing], and can clone a voice from as little as 10 seconds of audio. A shared voice marketplace exposes more than 2,000,000 public voices, so creators can either publish their own or pick a ready-made one instead of recording.
The platform is usage-based at its core. Subscription plans meter output in credits — the free tier includes 8,000 credits per month (about 7 minutes), while paid plans scale up to millions of credits — and the developer API is pure pay-as-you-go with no monthly minimum. TTS models (s2.1-pro, s2-pro, s1) are priced at $15 per million UTF-8 bytes, roughly 180,000 English words or about 12 hours of speech, with a free s2.1-pro-free model available under a fair-use policy. Speech-to-text and Voice Design are billed separately per audio hour and per request. Because non-Latin scripts use more bytes per character, they cost more per word than English.
Fish Audio suits builders and creators who want cheap, flexible voice synthesis and are comfortable trading some polish for lower cost and open models. The free tier and free API model make it easy to prototype, while the per-million-byte API pricing keeps high-volume generation predictable. For a more mature, feature-broad alternative, teams often compare it against ElevenLabs.
Starting price: $11/mo · Free tier: yes · Model: freemium
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Free | $0/mo | 8,000 credits/month (~7 minutes) · 500 characters per generation · 3 public voice slots · Instant voice cloning · Personal use |
| Plus | $11/mo | 250,000 credits/month (~200 minutes) · 15,000 characters per generation · 10 private + 1 professional voice slot · Voice Design feature · Commercial use |
| Pro | $75/mo | 2,000,000 credits/month (~1,620 minutes) · 30,000 characters per generation · 5 professional voice slots · 3 team seats · 7-day money-back guarantee |
| Max | $749/mo | 25,000,000 credits/month (~6,250 minutes) · 15 professional voice slots · 10 team seats · All Pro features |
| Enterprise | Custom | Pay-as-you-go with organizational controls · Zero data retention option · On-premise deployment · SOC2 compliance |
| Pros | Cons |
|---|---|
| Free s2.1-pro-free API model and an 8,000-credit monthly free tier make it cheap to start | Voice cloning raises consent and misuse concerns — cloning a real person's voice without permission is an ethical and legal risk |
| Pay-as-you-go API at $15 per million UTF-8 bytes (~12 hours of speech) with no monthly minimum | Free tier is limited (8,000 credits, ~7 minutes/month, 500 characters per generation) and is intended for personal use |
| Instant cloning needs only about 10 seconds of audio, a lower bar than many competitors | Output quality and voice consistency can trail more established platforms like ElevenLabs on some voices |
| Core Fish Speech models are open source, so self-hosting and inspection are possible | Non-Latin scripts such as Chinese, Japanese, and Korean cost more per word because they use multiple UTF-8 bytes per character |
AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages
AI voice platform for voice cloning, text-to-speech, and real-time voice agents, plus deepfake detection and audio watermarking
AI text-to-speech studio with 200+ voices across 35+ languages for voiceovers, dubbing, and voice agents
Low-latency voice AI platform behind the Sonic text-to-speech model, built for real-time agents
Text-to-speech and voice AI platform that reads any document aloud in 1,000+ voices
Yes. Fish Audio has a free tier with 8,000 credits per month (about 7 minutes of generation), a 500-character-per-generation cap, and instant voice cloning. The free plan is intended for personal use; commercial rights come with the paid plans starting at $11/month. Developers can also call the s2.1-pro-free API model at no cost under a fair-use policy.
The TTS API is pay-as-you-go with no subscription or monthly minimum. The s2.1-pro, s2-pro, and s1 models are priced at $15 per million UTF-8 bytes, which is roughly 180,000 English words or about 12 hours of speech. Speech-to-text (transcribe-1) is $0.36 per audio hour, and Voice Design is $0.01 per successful request. The s2.1-pro-free model is free under a fair-use policy.
Fish Audio can clone a voice from as little as 10 seconds of audio, a lower barrier than many competitors that require one to several minutes. Cloned voices use the same TTS endpoint and per-million-byte pricing as catalog voices.
Paid plans are Plus at $11/month (250,000 credits, ~200 minutes), Pro at $75/month (2,000,000 credits, ~1,620 minutes, 3 team seats), and Max at $749/month (25,000,000 credits, ~6,250 minutes, 10 team seats). Enterprise pricing is custom and billed annually. Annual billing and periodic promotions can lower the effective rate.
Fish Audio supports text-to-speech in 30+ languages, including English, Chinese, Japanese, Korean, French, German, Spanish, and Arabic. Note that non-Latin scripts consume more UTF-8 bytes per character, so they cost more per word on the API.
Partly. Fish Audio maintains open-source models — Fish Speech and Fish Diffusion — released through its OpenAudio research effort, so the underlying models can be self-hosted. The hosted platform, S2.1 Pro engine, voice marketplace, and API are commercial products.
Commercial use is included on the paid plans (Plus, Pro, Max, and Enterprise). The free tier is intended for personal use, so publishing or monetizing generated audio generally requires upgrading to a paid plan.