Skip to content
API Rates Per-unit API pricing — verified, dated, and logged when it changes
LLM inference APIs
DEEPSEEK, MOONSHOT AI, QWEN (ALIBABA)

DeepSeek vs Kimi K2 vs Qwen: cheapest capable LLM API

Published Pricing verified

Qwen’s qwen3.7‑flash undercuts the competition so aggressively that the other two models in this comparison feel like they are charging for a different tier of service. At $0.03 for input and $0.13 for output per 1M tokens, Alibaba’s model is literally four times cheaper than DeepSeek’s equivalent flash tier on input, and nearly a fifth of the cost on output. If your primary constraint is raw cost per token without sacrificing the ability to handle complex reasoning or long contexts, the math is not close. You are not picking between three similar options here; you are picking between the absolute floor of the market and two models that cost significantly more to run.

DeepSeek enters the conversation with deepseek‑v4‑flash, which lists an input cost of $0.065 and an output cost of $0.18 per 1M tokens. It also offers a cached input rate of $0.016. While this is respectable and significantly cheaper than standard API providers, it sits squarely in the middle of this specific trio. Moonshot AI’s kimi‑k2.5 is the most expensive of the three, demanding $0.45 for input and $2.25 for output per 1M tokens. That is fifteen times the input price of Qwen’s flash model and roughly seventeen times the output price. Unless you have a specific dependency on Moonshot’s ecosystem or require features explicitly tied to the Kimi architecture, it is difficult to justify the premium when a cheaper alternative is available for general-purpose tasks.

### The cost of capability
The pricing table below clarifies the disparity. Note that these are the specific model variants selected for this comparison to represent the "capable" tier of each vendor’s lineup that is also cost-competitive. Qwen’s qwen3.7‑flash is the clear outlier on the low end. DeepSeek’s deepseek‑v4‑flash offers a cached input discount that can help if your prompts are repetitive, with a rate of $0.016 per 1M cached tokens. Moonshot’s kimi‑k2.5 also offers a cached input rate of $0.07 per 1M tokens, but even with that discount, the base rates remain high. The output cost is where the bill really hurts for high-generation workloads, and here Qwen’s $0.13 per 1M tokens looks trivial next to Kimi’s $2.25.

VendorModelInputCached inputOutput
Qwenqwen3.7‑flash$0.03 per 1M tokens$0.006 per 1M tokens$0.13 per 1M tokens
DeepSeekdeepseek‑v4‑flash$0.065 per 1M tokens$0.016 per 1M tokens$0.18 per 1M tokens
Moonshot AIkimi‑k2.5$0.45 per 1M tokens$0.07 per 1M tokens$2.25 per 1M tokens

### Who should pick which
If you are building a high-volume application where every dollar counts—think customer support bots, large-scale data processing, or real-time chat interfaces—Qwen’s qwen3.7‑flash is the pragmatic choice. It offers the lowest entry point without forcing you into a proprietary ecosystem that might lock you in. DeepSeek is a strong second choice if you prefer its specific model architecture or if your workload benefits heavily from its cached input pricing; the $0.016 cached rate is competitive and can shave costs off repetitive prompt structures. Moonshot AI’s kimi‑k2.5 is the outlier in this comparison. It is priced for a different audience, likely those who have already integrated deeply into their platform or who require specific performance characteristics that justify the higher token spend. For a new project simply looking for the cheapest capable LLM API, the choice leans heavily toward Qwen, with DeepSeek as a reasonable fallback if you need the specific capabilities of the DeepSeek family.

Pricing verified on 2026-09-12.

Common questions

Which model is the cheapest per 1M tokens?

Qwen’s qwen3.7-flash is the cheapest, with input at $0.03 and output at $0.13 per 1M tokens.

Does DeepSeek offer a discounted rate for cached input?

Yes, DeepSeek’s deepseek-v4-flash charges $0.016 per 1M tokens for cached input.

How much more does Moonshot AI’s kimi-k2.5 cost compared to Qwen?

Moonshot AI’s kimi-k2.5 charges $0.45 per 1M tokens for input, whereas Qwen’s qwen3.7-flash charges $0.03 per 1M tokens for input.

Sources
  1. DeepSeek — pricing
  2. Moonshot AI — pricing
  3. Qwen (Alibaba) — pricing