Replicate
Run and fine-tune thousands of open-source AI models with one line of code via a cloud API, billed per second of GPU or CPU compute
AI cloud for running, fine-tuning, and deploying open-source models via serverless inference and on-demand GPU clusters
Together AI is an AI-native cloud for running open-source models like Llama, DeepSeek, and Qwen. It spans serverless per-token inference (Llama 3.3 70B at roughly $0.88 per 1M tokens), dedicated endpoints, fine-tuning, and on-demand GPU clusters (H100 from about $2.99/hour). With 200+ hosted models, free signup credits, and no monthly minimum, it suits developers who want open models in production without managing GPUs.
Together AI is an AI-native cloud built for open-source models. Rather than locking you into one provider’s proprietary model, it lets you run, fine-tune, and deploy open weights — Llama, DeepSeek, Qwen, Mistral, and 200+ others — across the full stack of infrastructure you might need. That ranges from serverless per-token inference, where you call a model through an API with no servers to manage, to dedicated endpoints for steady traffic, to raw GPU clusters of H100, H200, and B200 cards for training and large batch jobs.
Pricing is usage-based with no monthly minimum. Serverless inference is billed per token and varies by model — Llama 3.3 70B runs around $0.88 per 1M tokens — while GPU clusters are billed per GPU-hour, with H100 starting near $2.99/hour on reserved terms. Beyond text, the platform offers fine-tuning (LoRA and full), batch processing, multi-modal endpoints for images, audio, and embeddings, plus code sandboxes and object storage with zero egress fees. Free signup credits let teams evaluate before committing spend.
Together AI is a developer and team platform, not a consumer app. It fits builders who have chosen open models — for cost, control, customization, or data-residency reasons — and need somewhere to run them in production without operating their own GPUs.
Starting price: $0 · Free tier: yes · Model: usage-based
Price history tracked from June 2026
| Plan | Price | Includes |
|---|---|---|
| Serverless Inference | Pay per token | On-demand, no monthly minimum · Llama 3.3 70B around $0.88 per 1M tokens · 200+ open models available · Cached-token discounts |
| Dedicated Endpoints | From $6.49/hr | Reserved single-model capacity · Consistent low latency · Priced per GPU-hour (e.g. 1x H100 80GB) · For steady production traffic |
| GPU Clusters | From $2.99/hr | On-demand H100, H200, and B200 GPUs · Reserved terms lower the hourly rate · Scale to thousands of GPUs · For training and large workloads |
| Fine-Tuning | Pay per token | LoRA and full fine-tuning · Llama, Mistral, Qwen, DeepSeek and more · Priced per training token · No GPU management required |
| Pros | Cons |
|---|---|
| Broad open-model catalog (200+) with simple usage-based pricing and no monthly minimum | Per-token rates vary widely by model and change often, making cost forecasting harder |
| Covers the full stack — serverless inference, fine-tuning, and raw GPU clusters in one platform | GPU cluster and dedicated endpoint costs add up quickly for sustained workloads |
| Free signup credits and free/open models let you test before committing spend | Developer/API-focused with no consumer desktop or mobile app |
Run and fine-tune thousands of open-source AI models with one line of code via a cloud API, billed per second of GPU or CPU compute
Unified API to 400+ LLMs from 70+ providers through one OpenAI-compatible endpoint, with automatic failover and pass-through token pricing
Ultra-fast, low-cost inference for open models on custom LPU chips (groq.com — not xAI's Grok)
Free, open-source tool to run open-weight LLMs locally via CLI, desktop, or API — with an optional paid cloud for larger models
See all Together AI alternatives →
Together AI is a cloud platform for running, fine-tuning, and deploying open-source models such as Llama, DeepSeek, and Qwen. Developers use it to serve LLMs via API, fine-tune models on custom data, and rent GPU clusters for training without managing hardware.
Inference is usage-based, billed per token and varying by model — Llama 3.3 70B runs around $0.88 per 1M tokens. Dedicated endpoints and GPU clusters are billed per GPU-hour, with H100 starting near $2.99/hour on reserved terms. There is no monthly minimum.
Together AI provides free signup credits and access to free or open models so you can test before paying. After the credits, you pay only for what you use. The exact free credit amount can vary, so confirm it at signup.
Yes. Together AI offers LoRA and full fine-tuning across major open-model families including Llama, Mistral, Qwen, and DeepSeek. Fine-tuning is priced per training token, and the resulting model can be served from a dedicated endpoint.
Both serve open models via API, but Together AI is a broad platform spanning inference, fine-tuning, and GPU clusters, while Groq focuses on ultra-fast inference on its custom LPU chip. Together emphasizes breadth; Groq emphasizes raw speed.
Yes. Together AI rents on-demand and reserved GPU clusters using H100, H200, and B200 hardware, billed per GPU-hour. Reserved commitments lower the rate, and clusters can scale to thousands of GPUs for training or large-scale jobs.