Groq vs Together AI
Updated Jul 11, 2026 · pricing verified for both tools
Groq wins for raw speed — its LPU chips serve Llama 3.1 8B at roughly 840 tokens/second from $0.05 per 1M input tokens; Together AI wins on breadth with 200+ hosted models. Per-token prices are similar where their catalogs overlap, so the choice is speed versus scope: Groq for latency-critical chat, voice, and agent apps; Together AI for self-serve fine-tuning (LoRA and full), 405B-class models, and H100 GPU clusters from about $2.99/hour. Both expose OpenAI-compatible APIs and free entry points.
- Groq starts at
- $0
- Together AI starts at
- $0
- Groq free tier
- Yes
- Together AI free tier
- Yes
Both prices were verified against the official pricing pages on Jun 24, 2026.
Side by side
| Dimension | Groq | Together AI |
|---|---|---|
| Pricing | Usage-based; Llama 3.1 8B ~$0.05/1M input tokens; batch and cached input 50% off | Usage-based; Llama 3.3 70B ~$0.88/1M tokens; dedicated endpoints from $6.49/hr |
| Free tier | ✓ Genuine free tier with API key, no credit card; rate-limited (~30 requests/min) | Free signup credits (amount varies); free and open models for testing |
| Inference speed | ✓ Custom LPU silicon; hundreds to 1,000+ tokens/sec (Llama 3.1 8B ~840 tok/s) | GPU-based serving at standard industry throughput |
| Model catalog | Curated open models: Llama 3.x/4, GPT-OSS, Qwen3, Whisper, Orpheus TTS | ✓ 200+ open models including DeepSeek, Mistral, Qwen, and Llama 3.1 405B |
| Fine-tuning | Custom and fine-tuned models only via Enterprise sales | ✓ Self-serve LoRA and full fine-tuning, priced per training token |
| GPU infrastructure | No GPU rental — hosted inference only; on-prem via Enterprise | ✓ On-demand H100/H200/B200 clusters from ~$2.99/hr, scaling to thousands of GPUs |
| Platforms | Web console and OpenAI-compatible API | ✓ Web, OpenAI-compatible API, and CLI |
Pricing, verified
Groq
USAGE-BASED| Plan | Price |
|---|---|
| Free | Free |
| Developer | Pay per token |
| Enterprise | Custom |
Together AI
USAGE-BASED| Plan | Price |
|---|---|
| Serverless Inference | Pay per token |
| Dedicated Endpoints | From $6.49/hr |
| GPU Clusters | From $2.99/hr |
| Fine-Tuning | Pay per token |
When to pick Groq
- Latency is the product — voice agents, live chat, and real-time tools feel instant at LPU speeds of hundreds to 1,000+ tokens per second, a gap GPU-based providers don’t close.
- You want the cheapest path to testing: a genuine free tier issues an API key with no credit card, and the OpenAI-compatible endpoint means switching is roughly a two-line change.
- Your workload is batch-heavy — the Batch API and prompt caching each cut costs by 50%, which compounds on large summarization, classification, or enrichment jobs.
When to pick Together AI
- You need a model Groq doesn’t host — Together’s 200+ catalog spans DeepSeek, Mistral, Qwen, and heavyweights like Llama 3.1 405B, where Groq curates a shorter speed-optimized menu.
- Fine-tuning is on the roadmap: LoRA and full fine-tunes are self-serve and priced per training token, with the result deployable to a dedicated endpoint — on Groq this requires an Enterprise contract.
- You’ll eventually need raw compute — H100/H200/B200 clusters from about $2.99/hour let training, batch jobs, and inference live on one platform instead of a second vendor.
Bottom line
Groq and Together AI both serve open models through OpenAI-compatible APIs at similar per-token rates, so neither wins on price alone — they win on different axes. Groq is a specialist: its custom LPU hardware makes it the default for anything conversational or real-time, and its free tier plus batch discounts make it the cheapest sandbox. Together AI is a platform: the largest open-model catalog, self-serve fine-tuning, and GPU rental mean a team can prototype, customize, and scale without leaving. A common pattern is starting on Groq for speed and adding Together when fine-tuning or exotic models enter the picture — the shared API standard makes running both nearly free in engineering terms.
Frequently asked questions
Is Groq better than Together AI?
Groq wins for raw speed — its LPU chips serve Llama 3.1 8B at roughly 840 tokens/second from $0.05 per 1M input tokens; Together AI wins on breadth with 200+ hosted models. Per-token prices are similar where their catalogs overlap, so the choice is speed versus scope: Groq for latency-critical chat, voice, and agent apps; Together AI for self-serve fine-tuning (LoRA and full), 405B-class models, and H100 GPU clusters from about $2.99/hour. Both expose OpenAI-compatible APIs and free entry points. Updated Jul 11, 2026.
Is Groq cheaper than Together AI?
They start at the same price: $0 for Groq and $0 for Together AI. Both prices were verified against the official pricing pages on Jun 24, 2026.
Do Groq and Together AI have free tiers?
Yes — both do. Groq has a free tier (the Free plan). Together AI has a free tier.
Can I use Groq and Together AI on mobile or via an API?
Groq runs on the web and offers an API. Together AI runs on the web, the command line and offers an API.