Fish Audio
Voice cloning and text-to-speech platform with a shared voice marketplace and pay-as-you-go API, built on open-source Fish Speech models
Fish Audio is a voice-cloning and text-to-speech platform built on the open-source Fish Speech and OpenAudio models. The free tier gives 8,000 credits per month (about 7 minutes); paid plans start at $11/month, and the TTS API is pay-as-you-go at $15 per million UTF-8 bytes, roughly 12 hours of speech. Best for developers and creators who want low-cost, multilingual voice synthesis with instant cloning.
What is Fish Audio?
Fish Audio is an AI voice platform for text-to-speech and voice cloning, built on top of the open-source Fish Speech and Fish Diffusion models released through its OpenAudio research effort. Its hosted S2.1 Pro engine turns text into expressive speech in 30+ languages, supports emotion and effect tags like [angry], [whispering], and [laughing], and can clone a voice from as little as 10 seconds of audio. A shared voice marketplace exposes more than 2,000,000 public voices, so creators can either publish their own or pick a ready-made one instead of recording.
The platform is usage-based at its core. Subscription plans meter output in credits — the free tier includes 8,000 credits per month (about 7 minutes), while paid plans scale up to millions of credits — and the developer API is pure pay-as-you-go with no monthly minimum. TTS models (s2.1-pro, s2-pro, s1) are priced at $15 per million UTF-8 bytes, roughly 180,000 English words or about 12 hours of speech, with a free s2.1-pro-free model available under a fair-use policy. Speech-to-text and Voice Design are billed separately per audio hour and per request. Because non-Latin scripts use more bytes per character, they cost more per word than English.
Who is it for?
Fish Audio suits builders and creators who want cheap, flexible voice synthesis and are comfortable trading some polish for lower cost and open models. The free tier and free API model make it easy to prototype, while the per-million-byte API pricing keeps high-volume generation predictable. For a more mature, feature-broad alternative, teams often compare it against ElevenLabs.
- Developers integrating multilingual TTS or voice agents who want pay-as-you-go billing and a free model to prototype against before committing spend.
- Content creators and podcasters producing YouTube voiceovers, audiobooks to ACX/Audible specs, or character voices who want instant cloning from a short sample.
- Indie hackers and small teams who want to publish cloned voices to the marketplace or resell generated narration at a markup over the API rate.
- Technical users who value the open-source Fish Speech and Fish Diffusion models for self-hosting, inspection, or offline experimentation.
How much does Fish Audio cost?
Starting price: $11/mo · Free tier: yes · Model: freemium
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Free | $0/mo | 8,000 credits/month (~7 minutes) · 500 characters per generation · 3 public voice slots · Instant voice cloning · Personal use |
| Plus | $11/mo | 250,000 credits/month (~200 minutes) · 15,000 characters per generation · 10 private + 1 professional voice slot · Voice Design feature · Commercial use |
| Pro | $75/mo | 2,000,000 credits/month (~1,620 minutes) · 30,000 characters per generation · 5 professional voice slots · 3 team seats · 7-day money-back guarantee |
| Max | $749/mo | 25,000,000 credits/month (~6,250 minutes) · 15 professional voice slots · 10 team seats · All Pro features |
| Enterprise | Custom | Pay-as-you-go with organizational controls · Zero data retention option · On-premise deployment · SOC2 compliance |
What are Fish Audio's key features?
- Instant voice cloning from as little as 10 seconds of audio
- Text-to-speech in 30+ languages powered by the S2.1 Pro engine
- Emotion and effect tags such as [angry], [whispering], [laughing], and [pause]
- Shared voice marketplace with 2,000,000+ public voices
- Pay-as-you-go TTS API priced per million UTF-8 bytes, plus a free s2.1-pro-free model
- Speech-to-text transcription (transcribe-1) and Voice Design generation endpoints
- Open-source Fish Speech and Fish Diffusion models under the OpenAudio research effort
What people use Fish Audio for
- 01 Cloning a personal or brand voice from a short audio sample for consistent narration
- 02 Generating multilingual video voiceovers and YouTube narration across 30+ languages
- 03 Producing audiobook narration that meets ACX/Audible specifications
- 04 Adding low-latency spoken responses to chatbots and voice agents through the API
- 05 Browsing the 2,000,000+ voice marketplace to find a ready-made voice instead of recording
Pros and cons
| Pros | Cons |
|---|---|
| Free s2.1-pro-free API model and an 8,000-credit monthly free tier make it cheap to start | Voice cloning raises consent and misuse concerns — cloning a real person's voice without permission is an ethical and legal risk |
| Pay-as-you-go API at $15 per million UTF-8 bytes (~12 hours of speech) with no monthly minimum | Free tier is limited (8,000 credits, ~7 minutes/month, 500 characters per generation) and is intended for personal use |
| Instant cloning needs only about 10 seconds of audio, a lower bar than many competitors | Output quality and voice consistency can trail more established platforms like ElevenLabs on some voices |
| Core Fish Speech models are open source, so self-hosting and inspection are possible | Non-Latin scripts such as Chinese, Japanese, and Korean cost more per word because they use multiple UTF-8 bytes per character |
What are the best Fish Audio alternatives?
Cartesia is the closest Fish Audio alternative in our directory: it covers the same category (AI Audio & Music), starts at $0, has a free tier.
Ranked by category overlap with Fish Audio, then free-tier availability, then lowest verified starting price — computed from our verified data, never from sponsorships.
| Alternative | What it is | Starting price | Free tier | Price verified |
|---|---|---|---|---|
| Cartesia | Low-latency voice AI platform behind the Sonic text-to-speech model, built for real-time agents | $0 | yes | Jul 5, 2026 |
| ElevenLabs | AI voice platform for text-to-speech, voice cloning, and conversational agents in 70+ languages | $6/mo | yes | Jul 3, 2026 |
| Speechify | Text-to-speech and voice AI platform that reads any document aloud in 1,000+ voices | $11.58/mo | yes | Jun 11, 2026 |
| Murf AI | AI text-to-speech studio with 200+ voices across 35+ languages for voiceovers, dubbing, and voice agents | $19/mo | yes | Jun 11, 2026 |
| Resemble AI | AI voice platform for voice cloning, text-to-speech, and real-time voice agents, plus deepfake detection and audio watermarking | $0.0005/sec | no | Jun 30, 2026 |
How people make money with Fish Audio
- Publish custom cloned voices in the Fish Audio voice marketplace and earn per-use royalties each time other users generate speech with them
- Offer multilingual voiceover and audiobook production on Fiverr or ACX using instant cloning and 30+ language TTS, keeping input costs low with pay-as-you-go API billing
- Build a niche narration or chatbot-voice app on the API and resell generated audio to creators at a markup over the per-million-byte rate
Frequently asked questions
Is Fish Audio free?
Yes. Fish Audio has a free tier with 8,000 credits per month (about 7 minutes of generation), a 500-character-per-generation cap, and instant voice cloning. The free plan is intended for personal use; commercial rights come with the paid plans starting at $11/month. Developers can also call the s2.1-pro-free API model at no cost under a fair-use policy.
How much does the Fish Audio API cost?
The TTS API is pay-as-you-go with no subscription or monthly minimum. The s2.1-pro, s2-pro, and s1 models are priced at $15 per million UTF-8 bytes, which is roughly 180,000 English words or about 12 hours of speech. Speech-to-text (transcribe-1) is $0.36 per audio hour, and Voice Design is $0.01 per successful request. The s2.1-pro-free model is free under a fair-use policy.
How much audio does Fish Audio need to clone a voice?
Fish Audio can clone a voice from as little as 10 seconds of audio, a lower barrier than many competitors that require one to several minutes. Cloned voices use the same TTS endpoint and per-million-byte pricing as catalog voices.
How much does Fish Audio cost per month?
Paid plans are Plus at $11/month (250,000 credits, ~200 minutes), Pro at $75/month (2,000,000 credits, ~1,620 minutes, 3 team seats), and Max at $749/month (25,000,000 credits, ~6,250 minutes, 10 team seats). Enterprise pricing is custom and billed annually. Annual billing and periodic promotions can lower the effective rate.
What languages does Fish Audio support?
Fish Audio supports text-to-speech in 30+ languages, including English, Chinese, Japanese, Korean, French, German, Spanish, and Arabic. Note that non-Latin scripts consume more UTF-8 bytes per character, so they cost more per word on the API.
Is Fish Audio open source?
Partly. Fish Audio maintains open-source models — Fish Speech and Fish Diffusion — released through its OpenAudio research effort, so the underlying models can be self-hosted. The hosted platform, S2.1 Pro engine, voice marketplace, and API are commercial products.
Can I use Fish Audio voices commercially?
Commercial use is included on the paid plans (Plus, Pro, Max, and Enterprise). The free tier is intended for personal use, so publishing or monetizing generated audio generally requires upgrading to a paid plan.