Skip to content
AITrendTool

Fish Audio

Voice cloning and text-to-speech platform with a shared voice marketplace and pay-as-you-go API, built on open-source Fish Speech models

Fish Audio is a voice-cloning and text-to-speech platform built on the open-source Fish Speech and OpenAudio models. The free tier gives 8,000 credits per month (about 7 minutes); paid plans start at $11/month, and the TTS API is pay-as-you-go at $15 per million UTF-8 bytes, roughly 12 hours of speech. Best for developers and creators who want low-cost, multilingual voice synthesis with instant cloning.

Verified JUL 10, 2026 FREEMIUM Live
Screenshot of Fish Audio

What is Fish Audio?

Fish Audio is an AI voice platform for text-to-speech and voice cloning, built on top of the open-source Fish Speech and Fish Diffusion models released through its OpenAudio research effort. Its hosted S2.1 Pro engine turns text into expressive speech in 30+ languages, supports emotion and effect tags like [angry], [whispering], and [laughing], and can clone a voice from as little as 10 seconds of audio. A shared voice marketplace exposes more than 2,000,000 public voices, so creators can either publish their own or pick a ready-made one instead of recording.

The platform is usage-based at its core. Subscription plans meter output in credits — the free tier includes 8,000 credits per month (about 7 minutes), while paid plans scale up to millions of credits — and the developer API is pure pay-as-you-go with no monthly minimum. TTS models (s2.1-pro, s2-pro, s1) are priced at $15 per million UTF-8 bytes, roughly 180,000 English words or about 12 hours of speech, with a free s2.1-pro-free model available under a fair-use policy. Speech-to-text and Voice Design are billed separately per audio hour and per request. Because non-Latin scripts use more bytes per character, they cost more per word than English.

Who is it for?

Fish Audio suits builders and creators who want cheap, flexible voice synthesis and are comfortable trading some polish for lower cost and open models. The free tier and free API model make it easy to prototype, while the per-million-byte API pricing keeps high-volume generation predictable. For a more mature, feature-broad alternative, teams often compare it against ElevenLabs.

  • Developers integrating multilingual TTS or voice agents who want pay-as-you-go billing and a free model to prototype against before committing spend.
  • Content creators and podcasters producing YouTube voiceovers, audiobooks to ACX/Audible specs, or character voices who want instant cloning from a short sample.
  • Indie hackers and small teams who want to publish cloned voices to the marketplace or resell generated narration at a markup over the API rate.
  • Technical users who value the open-source Fish Speech and Fish Diffusion models for self-hosting, inspection, or offline experimentation.

How much does Fish Audio cost?

Starting price: $11/mo · Free tier: yes · Model: freemium

Pricing verified JUL 10, 2026

Price history tracked from June 2026

Fish Audio pricing tiers, verified against the official pricing page
Plan Price Includes
Free $0/mo 8,000 credits/month (~7 minutes) · 500 characters per generation · 3 public voice slots · Instant voice cloning · Personal use
Plus $11/mo 250,000 credits/month (~200 minutes) · 15,000 characters per generation · 10 private + 1 professional voice slot · Voice Design feature · Commercial use
Pro $75/mo 2,000,000 credits/month (~1,620 minutes) · 30,000 characters per generation · 5 professional voice slots · 3 team seats · 7-day money-back guarantee
Max $749/mo 25,000,000 credits/month (~6,250 minutes) · 15 professional voice slots · 10 team seats · All Pro features
Enterprise Custom Pay-as-you-go with organizational controls · Zero data retention option · On-premise deployment · SOC2 compliance

What are Fish Audio's key features?

  • Instant voice cloning from as little as 10 seconds of audio
  • Text-to-speech in 30+ languages powered by the S2.1 Pro engine
  • Emotion and effect tags such as [angry], [whispering], [laughing], and [pause]
  • Shared voice marketplace with 2,000,000+ public voices
  • Pay-as-you-go TTS API priced per million UTF-8 bytes, plus a free s2.1-pro-free model
  • Speech-to-text transcription (transcribe-1) and Voice Design generation endpoints
  • Open-source Fish Speech and Fish Diffusion models under the OpenAudio research effort

What people use Fish Audio for

  1. 01 Cloning a personal or brand voice from a short audio sample for consistent narration
  2. 02 Generating multilingual video voiceovers and YouTube narration across 30+ languages
  3. 03 Producing audiobook narration that meets ACX/Audible specifications
  4. 04 Adding low-latency spoken responses to chatbots and voice agents through the API
  5. 05 Browsing the 2,000,000+ voice marketplace to find a ready-made voice instead of recording

Pros and cons

Pros and cons of Fish Audio
Pros Cons
Free s2.1-pro-free API model and an 8,000-credit monthly free tier make it cheap to start Voice cloning raises consent and misuse concerns — cloning a real person's voice without permission is an ethical and legal risk
Pay-as-you-go API at $15 per million UTF-8 bytes (~12 hours of speech) with no monthly minimum Free tier is limited (8,000 credits, ~7 minutes/month, 500 characters per generation) and is intended for personal use
Instant cloning needs only about 10 seconds of audio, a lower bar than many competitors Output quality and voice consistency can trail more established platforms like ElevenLabs on some voices
Core Fish Speech models are open source, so self-hosting and inspection are possible Non-Latin scripts such as Chinese, Japanese, and Korean cost more per word because they use multiple UTF-8 bytes per character

What are the best Fish Audio alternatives?

See all Fish Audio alternatives →

How people make money with Fish Audio

  • Publish custom cloned voices in the Fish Audio voice marketplace and earn per-use royalties each time other users generate speech with them
  • Offer multilingual voiceover and audiobook production on Fiverr or ACX using instant cloning and 30+ language TTS, keeping input costs low with pay-as-you-go API billing
  • Build a niche narration or chatbot-voice app on the API and resell generated audio to creators at a markup over the per-million-byte rate

Frequently asked questions

Is Fish Audio free?

Yes. Fish Audio has a free tier with 8,000 credits per month (about 7 minutes of generation), a 500-character-per-generation cap, and instant voice cloning. The free plan is intended for personal use; commercial rights come with the paid plans starting at $11/month. Developers can also call the s2.1-pro-free API model at no cost under a fair-use policy.

How much does the Fish Audio API cost?

The TTS API is pay-as-you-go with no subscription or monthly minimum. The s2.1-pro, s2-pro, and s1 models are priced at $15 per million UTF-8 bytes, which is roughly 180,000 English words or about 12 hours of speech. Speech-to-text (transcribe-1) is $0.36 per audio hour, and Voice Design is $0.01 per successful request. The s2.1-pro-free model is free under a fair-use policy.

How much audio does Fish Audio need to clone a voice?

Fish Audio can clone a voice from as little as 10 seconds of audio, a lower barrier than many competitors that require one to several minutes. Cloned voices use the same TTS endpoint and per-million-byte pricing as catalog voices.

How much does Fish Audio cost per month?

Paid plans are Plus at $11/month (250,000 credits, ~200 minutes), Pro at $75/month (2,000,000 credits, ~1,620 minutes, 3 team seats), and Max at $749/month (25,000,000 credits, ~6,250 minutes, 10 team seats). Enterprise pricing is custom and billed annually. Annual billing and periodic promotions can lower the effective rate.

What languages does Fish Audio support?

Fish Audio supports text-to-speech in 30+ languages, including English, Chinese, Japanese, Korean, French, German, Spanish, and Arabic. Note that non-Latin scripts consume more UTF-8 bytes per character, so they cost more per word on the API.

Is Fish Audio open source?

Partly. Fish Audio maintains open-source models — Fish Speech and Fish Diffusion — released through its OpenAudio research effort, so the underlying models can be self-hosted. The hosted platform, S2.1 Pro engine, voice marketplace, and API are commercial products.

Can I use Fish Audio voices commercially?

Commercial use is included on the paid plans (Plus, Pro, Max, and Enterprise). The free tier is intended for personal use, so publishing or monetizing generated audio generally requires upgrading to a paid plan.