Skip to content
API Rates Per-unit API pricing — verified, dated, and logged when it changes
LLM inference APIs
QWEN (ALIBABA), MOONSHOT AI, DEEPSEEK

DeepSeek Flash Beats Qwen and Moonshot on Pure Output Cost

Published Pricing verified

When you need to pump out massive text, the price of each output token becomes the decisive factor. DeepSeek’s deepseek‑v4‑flash charges $0.1336 per 1M output tokens, just a hair above Qwen’s entry‑level qwen3.7‑flash at $0.13, and dramatically lower than Moonshot’s premium kimi‑k3 at $11.5502. That raw difference alone reshapes the economics of any high‑throughput generation pipeline.

### How input and cache fees shape the total bill
Both DeepSeek and Qwen keep input fees modest. DeepSeek’s base flash model costs $0.0668 per 1M input tokens and offers a cached‑input rate of $0.0134, while Qwen’s qwen3.7‑flash asks $0.03 for input and $0.006 for cached input. Moonshot’s cheapest Kimi tier, kimi‑k2.5, starts at $0.45 input with a $0.07 cache discount, but its output price of $2.25 per 1M tokens already eclipses the flash competitors. When a workload reuses system prompts, the cached‑input discount matters, yet the overwhelming driver remains output cost. DeepSeek’s flash line also includes a vision‑enhanced variant, deepseek‑v4‑flash‑vision‑exp, at $0.22 input, $0.007 cached input, and $0.66 output, still far cheaper than Moonshot’s top‑end rates.

### Matching models to budget and capability
If the priority is sheer cost efficiency for large‑scale text generation, DeepSeek’s flash family wins hands‑down. Its combination of low input, modest cache discount, and sub‑$0.14 output pricing keeps the total spend minimal even when prompts are repeated. Qwen’s qwen3‑next‑80b‑a3b‑instruct offers 80 billion parameters at $1.1 per 1M output tokens, a reasonable trade‑off for applications that demand higher model capacity without the premium Moonshot pricing. Moonshot’s kimi‑k3 delivers the most advanced reasoning and safety features, justified only for niche use‑cases where those capabilities outweigh the $2.25+ output cost and the higher input fees.

VendorModelInputCached inputOutput
Qwen (Alibaba)qwen3.7‑flash$0.03 per 1M tokens$0.006 per 1M tokens$0.13 per 1M tokens
Moonshot AIkimi‑k2.5$0.45 per 1M tokens$0.07 per 1M tokens$2.25 per 1M tokens
DeepSeekdeepseek‑v4‑flash$0.0668 per 1M tokens$0.0134 per 1M tokens$0.1336 per 1M tokens

For developers whose main goal is cheap generation at scale, DeepSeek’s flash tier is the clear choice. Teams that need Alibaba’s ecosystem or a Chinese‑language‑optimized model will find Qwen’s flash offerings a balanced middle ground. When the application calls for the most sophisticated reasoning and you can absorb premium rates, Moonshot’s Kimi line provides the specialized performance that justifies its price.

Pricing verified on 2026-09-12.

Common questions

Is there a free tier for any of these providers?

The current pricing tables list only paid rates; no free tier is mentioned.

Are the vendors billing per seat or per usage?

All three vendors bill strictly per token usage, with separate rates for input, cached input, and output.

What is the cheapest output price among the listed models?

DeepSeek’s **deepseek-v4-flash** charges $0.1336 per 1M output tokens, the lowest output rate shown.

Sources
  1. Qwen (Alibaba) — pricing
  2. Moonshot AI — pricing
  3. DeepSeek — pricing