When a business needs to churn out long responses, the price per output token can make a billion‑dollar difference. DeepSeek’s deepseek‑chat charges $1.0287 per 1M tokens, while OpenAI’s gpt‑4o‑mini sits at $0.6 and Google Gemini’s gemini‑3.1‑flash‑lite‑image at $1.5. For a workload that generates twice as many output tokens as input, the higher output fee of Gemini means the total cost climbs faster than either DeepSeek or OpenAI.
Input and cached‑input rates also shape the economics. DeepSeek’s deepseek‑chat input fee is $0.2574 per 1M tokens, with no cached‑input price listed, whereas gpt‑4o‑mini charges $0.15 per 1M input and $0.075 for cached input. Gemini’s gemini‑3.1‑flash‑lite starts at $0.25 per 1M input and offers a $0.025 cached discount. When a pipeline relies heavily on reusing prompts, Gemini’s cache discount can offset its higher output price, but for pure generation it lags.
The table below pulls the raw numbers straight from the vendors:
| Vendor | Model | Input | Cached input | Output |
|---|---|---|---|---|
| DeepSeek | deepseek‑chat | $0.2574 per 1M | — | $1.0287 per 1M |
| OpenAI | gpt‑4o‑mini | $0.15 per 1M | $0.075 per 1M | $0.6 per 1M |
| Google Gemini | gemini‑3.1‑flash‑lite | $0.25 per 1M | $0.025 per 1M | $1.5 per 1M |
If your priority is keeping output costs low while still enjoying a modern model, DeepSeek’s deepseek‑chat offers the cheapest generation rate, even though its input fee is higher. For use cases that involve heavy prompt reuse, OpenAI’s gpt‑4o‑mini gives the best overall savings by combining a modest input price with a strong cached‑input discount. Gemini’s gemini‑3.1‑flash‑lite‑image is the most expensive on output, making it suitable only when image‑enabled chat is a must and cost is secondary.
Pricing verified on 2026-09-12.