Skip to content
API Rates Per-unit API pricing — verified, dated, and logged when it changes
LLM inference APIs
OPENAI, GOOGLE GEMINI, ANTHROPIC

GPT-4o mini vs Gemini Flash Lite vs Claude 3 Haiku: the legacy budget tier

Published Pricing verified

The notion that the 'budget' tier of large language models is a small, manageable expense is increasingly outdated, but the absolute floor has not moved. GPT-4o mini, Gemini Flash Lite, and Claude 3 Haiku represent a specific historical snapshot of low-cost inference, yet their current price tags reveal a widening divergence that challenges the idea of a flat, affordable entry point. While these models were once grouped together as the go-to options for high-volume, low-latency tasks, their 2026 pricing structures reflect different strategic positions within their respective vendor lineups.

The divergence in input and output costs

The raw token rates for these three 'legacy' budget models highlight a significant spread in both input and output pricing. GPT-4o mini from OpenAI is listed at $0.15 per 1M input tokens, $0.075 per 1M cached input tokens, and $0.6 per 1M output tokens. Google Gemini 2.5 Flash Lite, which serves as the closest comparable in the current Gemini lineup, is priced at $0.1 per 1M input tokens, $0.01 per 1M cached input tokens, and $0.4 per 1M output tokens. Anthropic’s Claude 3 Haiku, the oldest of the three in terms of its specific versioning, sits at $0.25 per 1M input tokens, $0.03 per 1M cached input tokens, and $1.25 per 1M output tokens. The output cost gap is the most striking differentiator here, with Claude 3 Haiku’s output price more than double that of GPT-4o mini and over three times the rate of Gemini 2.5 Flash Lite. Input costs are more tightly clustered, ranging from $0.1 to $0.25 per million tokens, but the cached input rates reveal a major advantage for Google, where cached reads cost only $0.01 per 1M tokens compared to $0.075 for OpenAI and $0.03 for Anthropic.

VendorModelInputCached inputOutput
OpenAIgpt-4o-mini$0.15 per 1M tokens$0.075 per 1M tokens$0.6 per 1M tokens
Google Geminigemini-2.5-flash-lite$0.1 per 1M tokens$0.01 per 1M tokens$0.4 per 1M tokens
Anthropicclaude-3-haiku$0.25 per 1M tokens$0.03 per 1M tokens$1.25 per 1M tokens

Why the 'legacy' label matters for procurement

Selecting one of these models in the current landscape requires understanding that they are no longer the cheapest option available. OpenAI now offers gpt-5-nano at $0.05 per 1M input and $0.4 per 1M output, which is cheaper on input than GPT-4o mini. Google gemini-2.5-flash-lite is still a strong contender, but the newer gemini-3.1-flash-lite is listed at $0.25 per 1M input and $1.5 per 1M output, signaling that the 'lite' tier itself is moving up the price ladder. Anthropic’s Claude 3 Haiku is the most expensive of the three budget models, with an output cost of $1.25 per 1M tokens that actually rivals some of their mid-tier Sonnet models in terms of per-token expense. For teams that have built their infrastructure specifically around these three models, the pricing remains stable and predictable, but for new implementations, the value proposition is clearer with the newer, smaller variants from these same vendors. The 'legacy budget tier' is less a category of affordability and more a category of established performance, where the price includes the reliability of a model that has been tested at scale over years.

If your workload is dominated by cached reads, Gemini 2.5 Flash Lite’s $0.01 per 1M cached input rate makes it the most cost-efficient choice for that specific pattern. If you require the specific output quality and guardrails of the GPT ecosystem and are willing to pay $0.6 per 1M output tokens, GPT-4o mini remains a solid, if not the cheapest, option. For those already invested in the Anthropic API who need a predictable, if premium, budget model, Claude 3 Haiku’s $1.25 per 1M output price is a significant step up from the other two, making it less competitive for pure cost-driven use cases. The choice ultimately depends on whether you are optimizing for the absolute lowest token cost, in which case you should look at the 'nano' models from OpenAI and the 'gemma' models from Google, or if you are optimizing for the specific capabilities and stability of these three established workhorses.

Pricing verified on 2026-09-12.

Common questions

Is there a free tier for these specific models?

The provided pricing data does not mention a free tier for GPT-4o mini, Gemini 2.5 Flash Lite, or Claude 3 Haiku.

Which of the three has the lowest input cost?

Gemini 2.5 Flash Lite has the lowest input cost at $0.1 per 1M tokens, followed by GPT-4o mini at $0.15 and Claude 3 Haiku at $0.25.

How does the cached input price compare for Claude 3 Haiku?

Claude 3 Haiku’s cached input is priced at $0.03 per 1M tokens, which is lower than its standard input price of $0.25 per 1M tokens.

Sources
  1. OpenAI — pricing
  2. Google Gemini — pricing
  3. Anthropic — pricing