Together AI
AI cloud for running, fine-tuning, and deploying open-source models via serverless inference and on-demand GPU clusters
High-speed inference and fine-tuning cloud for open-weight LLMs, image, and audio models via a usage-based API
Fireworks AI is a high-speed inference and fine-tuning cloud for open-weight models like DeepSeek, Llama, and Qwen, processing over 30 trillion tokens a day. Serverless inference is usage-based — roughly $0.20 per 1M tokens for 4-16B models and $0.90 for larger ones — while dedicated H100/H200 GPUs run $7.00/hour. New accounts get $1 in free credits. Best for developers serving open models in production without managing GPUs.
Fireworks AI is an inference and training cloud built for open-weight generative models. Instead of locking you into one vendor’s proprietary model, it lets you serve, fine-tune, and scale open weights — DeepSeek, Llama, Qwen, Mistral, and others — through a single usage-based API. The platform spans serverless per-token inference (you call a model with no servers to manage), dedicated on-demand GPU deployments for steady traffic, and reserved capacity for large workloads. It runs at production scale, processing more than 30 trillion tokens per day for customers including Cursor, Vercel, and Notion.
Pricing is usage-based with no monthly minimum. Serverless text and vision inference is billed per token by model size — roughly $0.10 per 1M tokens under 4B parameters, $0.20 for 4-16B, and $0.90 for models over 16B — with cached input and batch inference both discounted to 50% of the standard rate. Dedicated deployments are billed per GPU-hour, with H100 80GB and H200 141GB at $7.00/hour and B200 at $10.00/hour. Fine-tuning (LoRA, full-parameter, SFT, and DPO) is priced per million training tokens starting at $0.50. Beyond text, Fireworks serves image models like FLUX.1 and audio models like Whisper V3, and its OpenAI- and Anthropic-compatible APIs make migrating existing code straightforward.
Fireworks AI is a developer and team platform, not a consumer app. It fits builders who have chosen open models — for speed, cost, control, or data-residency reasons — and need a fast, low-latency place to run them in production without operating their own GPUs. Teams comparing options often weigh it against Together AI and Groq.
Starting price: $0.10/1M tokens · Free tier: yes · Model: usage-based
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Serverless Inference | Pay per token | Under 4B models ~$0.10 per 1M tokens · 4B-16B models ~$0.20 per 1M tokens · Over 16B models ~$0.90 per 1M tokens · Cached input and batch at 50% of price |
| Dedicated Deployments | From $7.00/hr | H100 80GB and H200 141GB at $7.00/hr · B200 180GB at $10.00/hr · B300 288GB at $12.00/hr · Reserved single-model GPU capacity |
| Fine-Tuning | From $0.50/1M tokens | LoRA SFT from $0.50 per 1M training tokens (up to 16B) · 16.1B-80B models from $3.00 per 1M tokens · Full-parameter and DPO training supported · Deploy the tuned model to serverless or dedicated |
| Free Credits | $1 | $1 in free credits for new accounts · OpenAI- and Anthropic-compatible APIs · No monthly minimum or seat fee |
| Pros | Cons |
|---|---|
| Fast, cost-efficient inference — customers like Notion report latency dropping from about 2 seconds to 350 milliseconds | Only $1 in free credits, so there is little room to evaluate at scale before paying |
| Simple size-based usage pricing (around $0.20 per 1M tokens for 4-16B models) with no monthly minimum | Per-token rates vary by model and headline models carry separate premium pricing, making cost forecasting harder |
| Covers the full stack — serverless inference, dedicated GPUs, and fine-tuning — in one platform | Developer- and API-focused with no consumer desktop or mobile app |
| OpenAI- and Anthropic-compatible APIs make switching from proprietary providers low-effort | Dedicated GPU hours at $7.00/hr and up add up quickly for sustained workloads |
AI cloud for running, fine-tuning, and deploying open-source models via serverless inference and on-demand GPU clusters
Ultra-fast, low-cost inference for open models on custom LPU chips (groq.com — not xAI's Grok)
Run and fine-tune thousands of open-source AI models with one line of code via a cloud API, billed per second of GPU or CPU compute
Usage-based inference cloud for generative media — 1,000+ image, video, and audio model APIs (FLUX, Kling, Veo) plus serverless GPUs from $1.89/hour
Unified API to 400+ LLMs from 70+ providers through one OpenAI-compatible endpoint, with automatic failover and pass-through token pricing
Fireworks AI is a cloud platform for running, fine-tuning, and deploying open-weight models such as DeepSeek, Llama, and Qwen. Developers use it to serve LLMs, image, and audio models via API, fine-tune them on custom data, and rent dedicated GPUs — without managing hardware. The platform processes over 30 trillion tokens per day.
Inference is usage-based and billed per token by model size: roughly $0.10 per 1M tokens under 4B, $0.20 for 4-16B, and $0.90 for models over 16B. Dedicated deployments are billed per GPU-hour — H100 80GB and H200 141GB at $7.00/hour, B200 at $10.00/hour. Cached input and batch inference are billed at 50% of the standard rate.
Fireworks AI gives new accounts $1 in free credits to start building. There is no ongoing free tier beyond that credit and no monthly minimum, so after the credit you pay only for the tokens and GPU hours you use.
Yes. Fireworks supports LoRA, full-parameter, SFT, and DPO fine-tuning priced per million training tokens — LoRA SFT starts at $0.50 per 1M tokens for models up to 16B and $3.00 for 16.1B-80B models. The tuned model can then be deployed to serverless or a dedicated endpoint.
Both are inference clouds for open models with per-token and per-GPU-hour billing. Fireworks emphasizes low-latency serving and its FireAttention-style optimizations, with customers like Cursor, Vercel, and Notion citing large speedups. Together AI leans on a very broad 200+ model catalog. Pricing and available models differ, so compare the specific model you need.
Yes. Beyond text LLMs, Fireworks serves image models like FLUX.1 and audio models like Whisper V3 Large through the same usage-based API, so a single account can cover text, vision, image generation, and transcription.
Yes. Fireworks exposes OpenAI- and Anthropic-compatible APIs, so applications built against those SDKs can often switch to Fireworks with minimal code changes by pointing at the Fireworks endpoint.