Skip to content
API Rates Per-unit API pricing — verified, dated, and logged when it changes
LLM inference APIs
ANTHROPIC, GOOGLE GEMINI

Gemini’s gemma‑3‑4b‑it slashes output costs compared to Claude‑3‑haiku, but Claude still leads on capability

Published Pricing verified

The invoice shock of $0.1 versus $1.25 per million output tokens reads like a headline: Gemini can churn cheap text while Anthropic charges a premium for the same token count.

### Where the dollars land per token
| Vendor | Model | Output per 1M tokens |
|--------|-------|----------------------|
| Anthropic | claude‑3‑haiku | $1.25 |
| Anthropic | claude‑sonnet‑5 | $10 |
| Anthropic | claude‑opus‑5 | $25 |
| Google Gemini | gemma‑3‑4b‑it | $0.1 |
| Google Gemini | gemma‑3‑12b‑it | $0.15 |
| Google Gemini | gemma‑3‑27b‑it | $0.45 |
| Google Gemini | gemma‑4‑26b‑a4b‑it | $0.22 |
| Google Gemini | gemma‑4‑31b‑it | $0.34 |
| Google Gemini | gemini‑3.1‑flash‑lite | $1.5 |
| Google Gemini | gemini‑3.5‑flash‑lite | $2.5 |

Gemini’s gemma‑3‑4b‑it delivers the lowest output price at $0.1, a ten‑fold advantage over Anthropic’s cheapest output rate. Even Gemini’s flash‑lite line, at $1.5, remains cheaper than Claude‑3‑haiku’s $1.25 output cost when you factor in the higher‑quality claims of Claude’s larger models. For workloads that are token‑hungry—batch summarization, log analysis, or large‑scale content generation—the savings multiply dramatically.

### Trade‑offs beyond the meter
Claude‑3‑haiku’s $1.25 output price is already a bargain for a model that consistently produces coherent, nuanced prose, but the jump to Claude‑sonnet‑5 ($10) and Claude‑opus‑5 ($25) reflects a steep premium for richer context windows and deeper reasoning. If your product depends on those advanced capabilities—complex code generation, multi‑turn reasoning, or high‑stakes copywriting—Anthropic’s higher rates may be justified. Gemini’s gemma family excels at raw throughput; its flash‑lite variants add a modest boost in speed and multimodal support at still‑reasonable prices, but they lack the depth of Claude‑sonnet and Claude‑opus.

For startups racing to ship features on a tight burn rate, Gemini’s gemma‑3‑4b‑it offers the most cost‑effective path to scale. Enterprises that need the highest quality, longest context, or specialized safety tuning should budget for Claude‑sonnet‑5 or Claude‑opus‑5 despite the higher per‑token price. In short, choose Gemini when volume and budget dominate; choose Anthropic when quality and capability outweigh raw cost.

Pricing verified on 2026-09-12.

Common questions

Is there a free tier for either vendor?

No free tier is listed in the current pricing for either Anthropic or Google Gemini.

Are the prices billed per seat or per usage?

The prices are per 1M tokens, so they are billed per usage.

What is the cheapest paid plan cost?

The cheapest output cost is Google Gemini gemma‑3‑4b‑it at $0.1 per 1M tokens.

Sources
  1. Anthropic — pricing
  2. Google Gemini — pricing