Replicate
Run and fine-tune thousands of open-source AI models with one line of code via a cloud API, billed per second of GPU or CPU compute
Usage-based inference cloud for generative media — 1,000+ image, video, and audio model APIs (FLUX, Kling, Veo) plus serverless GPUs from $1.89/hour
fal is a usage-based inference cloud for generative media: 1,000+ image, video, audio, and 3D model APIs plus serverless GPUs. There is no subscription — images run roughly $0.02–$0.04 each (FLUX Kontext Pro $0.04/image), video costs $0.05–$0.40 per second (Veo 3 at the top end), and H100 GPUs start at $1.89/hour. You pay only for successful outputs. Best for developers building media features into apps; there is no consumer-facing product.
fal is an inference cloud built specifically for generative media. Instead of training or publishing its own frontier models, it hosts more than 1,000 production-ready image, video, audio, and 3D models — FLUX 2, Seedream V4, Kling 3.0, Veo 3.1, Wan 2.5, Seedance 2.0 — behind one API, one SDK, and one bill. Pricing is output-based: an image costs a fixed few cents (FLUX Kontext Pro is $0.04), video is metered per second ($0.05 for Wan 2.5 up to $0.40 for Veo 3), and failed generations or queue time cost nothing. The company’s proprietary inference engine is the core pitch — it claims up to 10x faster generation than stock serving stacks, with a 99.99% uptime guarantee.
Beyond hosted model APIs, fal sells raw serverless GPU compute for custom or fine-tuned models: a discounted H100 runs $1.89/hour, an H200 $2.10/hour, and a B200 $3.49/hour, all billed per second with no idle cost. Every endpoint gets a browser playground, and the Workflows tool chains multiple models into a single pipeline — generate an image, then animate it — without glue code. Enterprise customers get dedicated H100/H200/B200 clusters, SOC 2 compliance, SSO, private endpoints, and custom per-endpoint pricing. fal reports over 1,500,000 developers on the platform, and it is the inference layer behind products from Canva, Perplexity, Poe, and PlayAI.
fal is for people who ship software, not people who want to type prompts into a finished app. There is no consumer product — the value is that a developer can add frontier image or video generation to an application in an afternoon, pay only for successful outputs, and swap models without changing vendors. If you just want to make pictures, a wrapper product built on fal will serve you better than fal itself.
Starting price: $0.02/megapixel · Free tier: no · Model: usage-based
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Image models | $0.02–$0.06/image | Seedream V4 — $0.03 per image · FLUX Kontext Pro — $0.04 per image · Nano Banana — $0.0398 per image · Qwen Image — $0.02 per megapixel · Higher resolutions priced proportionally |
| Video models | $0.05–$0.40/second | Wan 2.5 — $0.05 per second · Kling 2.5 Turbo Pro — $0.07 per second · Veo 3 — $0.40 per second · Ovi — $0.20 flat per video |
| Serverless GPUs | From $1.89/h | H100 80GB — $1.89/h (list $3.99/h) · H200 141GB — $2.10/h (list $4.50/h) · B200 180GB — $3.49/h (list $6.25/h) · B300 288GB — $4.49/h (list $8.50/h) · RTX PRO 6000 96GB — $1.10/h (list $2.99/h) · Billed per second of compute |
| Enterprise | Custom | Custom per-endpoint pricing and volume discounts · Dedicated compute clusters (H100, H200, B200 VMs) · SOC 2 compliance, SSO, private endpoints · Usage analytics and 24/7 priority support |
| Pros | Cons |
|---|---|
| You pay only for successful outputs — server errors and queue time are free, which is rare among inference clouds | Developer-only platform — there is no consumer app, so non-technical users need a wrapper product built on top of it |
| One API surface covers frontier image, video, and audio models the same week they launch, instead of one vendor per model | Costs scale linearly with volume and can spike fast: Veo 3 at $0.40/second is $24 per minute of generated video |
| Discounted H100 at $1.89/hour undercuts most on-demand GPU clouds; per-second billing means no idle spend | No permanent free tier — signup credits are reported at around $20, and free credits expire |
| Every endpoint has a browser playground, so you can validate quality and cost before writing integration code | With 1,000+ endpoints each carrying its own rate, budgeting requires checking per-model pricing; upstream model providers can reprice or deprecate endpoints |
Run and fine-tune thousands of open-source AI models with one line of code via a cloud API, billed per second of GPU or CPU compute
AI cloud for running, fine-tuning, and deploying open-source models via serverless inference and on-demand GPU clusters
API-first AI image generator from Black Forest Labs with per-image pay-as-you-go pricing
AI image suite aggregating 40+ models like FLUX and Stable Diffusion for text-to-image, editing, and video, plus a separate developer API
fal is pay-per-use with no subscription. Most models bill per output: images at roughly $0.02–$0.06 each (or per megapixel, scaling with resolution), video at $0.05–$0.40 per second, and some models at a flat rate per generation. Models without a fixed output price fall back to per-second billing on the GPU that ran the request.
No permanent free plan exists. Third-party trackers report around $20 in signup credits for new accounts, but fal's official pricing page does not advertise a free tier, and free credits expire. After credits, everything is metered pay-per-use.
Both are usage-based model-hosting clouds, but fal specializes in generative media (image, video, audio, 3D) and pushes latency hard — its proprietary inference engine claims up to 10x faster generation. Replicate hosts a broader general-purpose catalog including language models. fal also sells raw serverless GPU time from $1.89/hour for custom models.
Over 1,000 production-ready models, including FLUX 2 and Seedream V4 for images, Kling 3.0, Veo 3.1, Wan 2.5, and Seedance 2.0 for video, plus audio and 3D generation models. New frontier media models typically appear on fal within days of release.
Yes. Serverless GPUs let you deploy custom or fine-tuned models with per-second billing: a discounted H100 (80GB) costs $1.89/hour, H200 (141GB) $2.10/hour, and B200 (180GB) $3.49/hour. Enterprise customers can get dedicated compute clusters.
No. Billing is output-based — you are charged only for successful outputs, never for server errors or time a request spends waiting in the queue.
fal states it serves over 1,500,000 developers, with Canva, Perplexity, Poe, and PlayAI named as customers. It positions itself as the inference layer behind consumer AI apps rather than a consumer product itself.