Skip to content
AITrendTool

Baseten

Usage-based AI inference cloud to deploy and serve ML models on dedicated GPUs with per-minute billing, per-token Model APIs, and autoscaling

Baseten is a usage-based AI inference cloud for deploying and serving ML models on dedicated GPUs. Pricing is pay-as-you-go with a $0 base plan: GPUs bill per minute from $0.631/hr (T4) to $6.50/hr (H100) and $9.98/hr (B200), plus per-token Model APIs and scale-to-zero autoscaling. Best for teams needing production-grade inference with low cold starts.

Verified JUL 17, 2026 USAGE-BASED Live
Screenshot of Baseten

What is Baseten?

Baseten is a production inference cloud for deploying and serving machine-learning models. Instead of renting bare GPUs and wiring up your own serving layer, you package a model — open-source, fine-tuned, or fully custom — with Baseten’s Truss framework and a command-line tool, then serve it behind an autoscaling REST API endpoint. Dedicated deployments run on named GPU types billed by the minute, ranging from a T4 at about $0.631 per hour up to an H100 at $6.50 per hour and a B200 at $9.98 per hour, with autoscaling that can scale to zero so you only pay for compute that is actively serving traffic.

Alongside dedicated deployments, Baseten offers Model APIs: pre-optimized, multi-tenant endpoints for popular open-weight models such as DeepSeek, Kimi, GLM, and GPT-OSS, billed per token (roughly $0.10 to $4.40 per million tokens depending on the model) so you can prototype without provisioning your own hardware. The platform’s differentiator is the Baseten Inference Stack — custom GPU kernels, advanced decoding techniques, and caching aimed at higher throughput and fast cold starts. It supports image generation, transcription, text-to-speech, embeddings, and compound multi-model chains, targets 99.99% uptime, and is SOC 2 Type II and HIPAA compliant, with self-host and hybrid options for regulated workloads.

Who is it for?

Baseten is aimed at engineering and ML teams that have moved past prototyping and need production-grade inference they do not want to build and operate themselves. The $0 pay-as-you-go Basic plan lowers the barrier to entry, but the real value shows up when you are serving real traffic at scale and care about cold-start latency, throughput, and reliability rather than the absolute cheapest GPU-hour. Budget-focused hobbyists are usually better served by cheaper raw-GPU options like Modal or RunPod.

  • ML and platform engineers who need to deploy custom or fine-tuned models as autoscaling endpoints without managing Kubernetes, GPU provisioning, or serving infrastructure themselves.
  • AI product teams building agents, automation backends, or compound pipelines that require low-cold-start, high-throughput inference with predictable per-minute GPU pricing.
  • Developers prototyping with open-weight models who want per-token Model APIs first, then a clean path to dedicated deployments as usage grows.
  • Regulated and enterprise organizations that need SOC 2 Type II and HIPAA compliance, data residency control, custom SLAs, and self-host or hybrid deployment options.

How much does Baseten cost?

Starting price: $0 · Free tier: no · Model: usage-based

Pricing verified JUL 17, 2026

Price history tracked from June 2026

Baseten pricing tiers, verified against the official pricing page
Plan Price Includes
Basic $0/mo Pay-as-you-go — only active compute is billed · Dedicated deployments, Model APIs, and Training · GPUs per minute: T4 $0.631/hr, H100 $6.50/hr, B200 $9.98/hr · Fast cold starts and scale-to-zero autoscaling · SOC 2 Type II and HIPAA compliant
Pro Volume discounts Everything in Basic plus priority GPU access · Dedicated compute and higher rate limits · Engineering expertise for optimization · Negotiated volume pricing on GPU-hours
Enterprise Custom Everything in Pro plus custom SLAs · Self-host and hybrid deployment options · On-demand flex compute · Data residency control

What are Baseten's key features?

  • Dedicated deployments on GPUs from T4 ($0.631/hr) to H100 ($6.50/hr) and B200 ($9.98/hr), billed per minute
  • Model APIs — pre-optimized per-token endpoints for models like DeepSeek V4, Kimi, GLM, and GPT-OSS
  • Autoscaling with scale-to-zero, so only active compute is billed and idle GPUs cost nothing
  • Baseten Inference Stack with custom kernels, advanced decoding, and caching for higher throughput
  • Truss packaging framework plus a CLI for deploying custom, open-source, or fine-tuned models
  • Training on the same inference-optimized infrastructure at the same per-minute GPU rates
  • Cross-cloud deployment targeting 99.99% uptime, with SOC 2 Type II and HIPAA compliance

What people use Baseten for

  1. 01 Deploying open-source or fine-tuned LLMs as production inference endpoints with autoscaling
  2. 02 Serving image generation, transcription, text-to-speech, or embedding models behind a REST API
  3. 03 Prototyping with per-token Model APIs before committing to a dedicated GPU deployment
  4. 04 Running compound AI pipelines that chain multiple models with low latency
  5. 05 Powering AI agents and automation backends that need reliable, low-cold-start model inference

Pros and cons

Pros and cons of Baseten
Pros Cons
Transparent, published per-minute GPU rates with no idle-time charges and scale-to-zero autoscaling On-demand H100 at ~$6.50/hr is well above budget GPU clouds — Modal lists H100s near $3.95/hr and RunPod is lower still
Both per-token Model APIs and dedicated GPU deployments in one platform, so you can prototype then scale GPUs are billed per minute, coarser than the per-second billing offered by Modal and some rivals
Optimized inference stack (custom kernels, caching) targets higher throughput and fast cold starts No permanent free tier — only starter credits, with no fixed free-credit dollar amount published
Enterprise-grade compliance (SOC 2 Type II, HIPAA) plus self-host and hybrid deployment options Geared toward production teams; the pricing and inference stack are overkill for hobby or one-off experiments

What are the best Baseten alternatives?

See all Baseten alternatives →

Frequently asked questions

Is Baseten free?

Baseten has no permanent free tier. The Basic plan costs $0 per month and is pay-as-you-go, so you only pay for active compute. New accounts receive some starter credits to experiment with deployments, but Baseten does not publish a fixed free-credit dollar amount on its pricing page.

How much does an H100 cost on Baseten?

Baseten bills a dedicated H100 (80GB) GPU at $0.10833 per minute, which works out to about $6.50 per hour. A 40GB H100 MIG slice is $3.75 per hour, and the newer B200 (180GB) is $9.98 per hour. You are only billed while a replica is actively serving traffic.

Does Baseten charge for idle time?

No. Baseten only bills for active compute, and its autoscaling can scale deployments down to zero replicas when there is no traffic. That means you avoid paying for idle GPUs between requests, though a cold start applies when a scaled-to-zero deployment spins back up to serve the next request.

What are Baseten Model APIs?

Model APIs are pre-optimized, multi-tenant endpoints for popular open-weight models such as DeepSeek, Kimi, GLM, and GPT-OSS. They are billed per token rather than per GPU-hour, typically from about $0.10 to $4.40 per million tokens depending on the model. They let you prototype without managing dedicated infrastructure.

How does Baseten pricing compare to Modal or RunPod?

Baseten's on-demand H100 at about $6.50 per hour is higher than several rivals; Modal lists H100s near $3.95 per hour and bills per second, while Baseten bills GPUs per minute. You pay a premium for Baseten's optimized inference stack, fast cold starts, and forward-deployed engineering support rather than the cheapest raw GPU.

Does Baseten offer an API and CLI?

Yes. Baseten is API-first: you package models with its Truss framework and a command-line tool, then call them over REST API endpoints. A web dashboard handles deployments, logs, and autoscaling settings, but there is no desktop or mobile app for managing the platform.