Replicate
Run and fine-tune thousands of open-source AI models with one line of code via a cloud API, billed per second of GPU or CPU compute
5 tools ranked · last updated Jul 30, 2026 · how we picked
The best pay-as-you-go AI API in 2026 is Replicate, which runs thousands of open-source models behind one endpoint and bills per second of compute — from $0.000025 a second on CPU to about $5.49 an hour on an H100, with no monthly plan to sign. Groq is the cheapest for text at roughly $0.05 per million input tokens on Llama 3.1 8B, and it and Deepgram are the only two of these five with a permanent free tier.
Prices last verified Jul 30, 2026 against official pricing pages.
Replicate runs thousands of open-source models — image, video, audio, and language — through one API and bills per second of compute, which means an app with ten users a day and one with ten thousand pay on exactly the same terms. Hardware spans CPU at $0.000025 per second up to an Nvidia H100 at $0.001525 per second (about $5.49 an hour), and some public models bill per run instead, such as FLUX 1.1 Pro at $0.04 per image. Cog packaging, fine-tuning, and deployment cover the full custom-model lifecycle if you outgrow the public catalog. It fits developers shipping model-backed features without operating GPUs. The caveats: there is no permanent free tier beyond limited free runs on a curated set, and private models and deployments bill for all online time — including setup and idle — not just active inference.
fal is a usage-based inference cloud aimed squarely at generative media: 1,000+ image, video, audio, and 3D model endpoints, often carrying frontier models the same week they launch. There is no subscription at any tier. Images run roughly $0.02–$0.06 each (FLUX Kontext Pro is $0.04), video costs $0.05–$0.40 per second depending on the model, and serverless H100 GPUs start at $1.89 an hour, which undercuts most on-demand GPU clouds. Its best billing detail is that you pay only for successful outputs — server errors and queue time are free, which is rare in this category — and every endpoint has a browser playground for checking quality and cost before writing code. The caveats: costs scale linearly and can spike fast (Veo 3 at $0.40 a second is $24 per minute of video), and signup credits are reported at around $20 and expire.
Groq runs open models — Llama, GPT-OSS, Qwen — on its own LPU inference chips, and the result is both the fastest and among the cheapest text inference available: hundreds to over 1,000 tokens per second, with Llama 3.1 8B at about $0.05 per million input tokens at roughly 840 tokens a second. GroqCloud has a genuine free tier that needs no credit card, then pay-as-you-go per-token pricing with no plan to cancel, and the API is OpenAI-compatible so migrating an existing integration is a base-URL change. It fits developers who need cheap, fast inference for agents, classification, or high-volume summarisation. The caveats: it serves open and hosted models only — no Claude- or GPT-class frontier models — free-tier rate limits are tight for production, and per-token rates and the model lineup change over time. Note this is Groq, not xAI’s Grok.
Deepgram covers the voice side of the same billing model: speech-to-text, text-to-speech, and full voice agents through one API, billed by usage with no minimums and $200 in free credits on a new account — the largest starting allowance in this list. Nova-3 speech-to-text runs about $0.0077 per minute for pre-recorded audio, Aura-2 text-to-speech about $0.030 per 1,000 characters, and the Voice Agent API starts near $0.075 per minute, with streaming latency around 200–300ms that holds up for real-time conversation. Self-hosting, SOC 2, and on-premise options make it viable for regulated, high-volume workloads. The caveats: it is developer-only with no consumer app, multilingual and pre-recorded rates cost noticeably more than the headline monolingual streaming price, and its text-to-speech voices are less expressive than ElevenLabs for premium narration.
Retell AI builds voice agents that answer and place phone calls, transfer to humans, navigate IVR menus, and book appointments — and it is the rare enterprise-adjacent voice platform sold without an annual contract. Pricing is $0.07 to $0.31 per minute depending on the model and voice you pick, from $0.002 per message for chat agents, with $10 in free credits and 20 concurrent calls included on the self-serve plan. Component rates are published individually — voice infrastructure at $0.055 a minute, platform voices at $0.015, ElevenLabs voices at $0.040 — so a call’s cost can be calculated before you build it. The caveats: that $0.07 floor only applies with the cheapest model and voice, telephony is billed separately on top of the platform rate, and HIPAA coverage, SSO, and uncapped concurrency all require an Enterprise agreement with no published price.
We ranked these five on one criterion first — can you call the API and pay only for what you use, with no seat licence and no monthly minimum — then on published per-unit rates, breadth of what the endpoint covers, and how honestly the headline price predicts the real bill. Replicate leads on breadth and the granularity of per-second billing; Groq leads on raw cost per token and is one of only two here, with Deepgram, that carries a permanent free tier. Usage billing cuts both ways: three of these five have no free tier at all, and every one of them can produce a surprising invoice if a workload runs hotter than planned, so set spend alerts before launch. Prices were verified between June and July 2026 against each product’s official pricing page; no tool paid or was paid to appear.
Run and fine-tune thousands of open-source AI models with one line of code via a cloud API, billed per second of GPU or CPU compute
Usage-based inference cloud for generative media — 1,000+ image, video, and audio model APIs (FLUX, Kling, Veo) plus serverless GPUs from $1.89/hour
Ultra-fast, low-cost inference for open models on custom LPU chips (groq.com — not xAI's Grok)
Developer-first voice AI API for speech-to-text, text-to-speech, and real-time voice agents
Pay-as-you-go platform for building AI phone agents that handle calls, transfers, and bookings
Every tool in this list has a full profile in our directory with pricing verified against its official pricing page on the date shown on its stamp. Ranking reflects verified pricing, free-tier generosity, platform coverage, and documented capabilities — not sponsorships. Nobody can pay to appear here. Read the full methodology.
Yes — 2 of the 5 tools here have a free tier: Groq, Deepgram. Pricing verified Jul 30, 2026.
5 of the 5 tools list an API: Replicate, fal, Groq, Deepgram, Retell AI.