# The token‑meter shows Moonshot’s Kimi K2 at $0.57 per 1M input and $2.3 per 1M output, while DeepSeek’s flash tier starts at $0.065 per 1M input and $0.18 per 1M output, and Alibaba’s Qwen‑3.7‑flash charges $0.03 per 1M input and $0.13 per 1M output. The spread is enough to flip a buying decision in seconds.
## How the rates stack up
| Vendor | Model | Input | Cached input | Output |
|--------|-------|-------|--------------|--------|
| Moonshot AI | kimi‑k2 | $0.57 per 1M tokens | – | $2.3 per 1M tokens |
| DeepSeek | deepseek‑v4‑flash‑0731 | $0.065 per 1M tokens | $0.016 per 1M tokens | $0.18 per 1M tokens |
| Qwen (Alibaba) | qwen3.7‑flash | $0.03 per 1M tokens | $0.006 per 1M tokens | $0.13 per 1M tokens |
Moonshot’s Kimi line is the most expensive on the output side, even its cheapest variant (kimi‑k2.5) still asks $2.25 per 1M output tokens. DeepSeek offers a middle ground: its flash‑0731 tier is dramatically cheaper than Moonshot but still above Qwen’s flash offering. Qwen‑3.7‑flash undercuts both on every metric, with the lowest input, cached input, and output rates of the three. If you are sensitive to output cost – which most production workloads are – Qwen’s flash model gives the biggest savings per million tokens generated.
## Who each price tier serves
Moonshot’s pricing makes sense for teams that need the specific capabilities of the Kimi series – for example, the Kimi‑K2‑Thinking variant that adds a cached‑input discount for repeated prompts. Those teams are often willing to pay a premium for the model’s Chinese‑language fine‑tuning and latency guarantees. DeepSeek’s flash‑0731 model is aimed at developers who want a solid Chinese‑language model without the Moonshot premium; the cached‑input discount helps when you repeatedly query the same context, but the output price remains higher than Qwen’s flash tier. Qwen‑3.7‑flash is built for high‑volume applications – chatbots, content generation, or batch processing – where every cent counts. Its low cached‑input rate also favors workloads that reuse embeddings or system prompts.
In practice, the choice collapses to three questions: Do you need Moonshot’s specialized Kimi features? If not, is DeepSeek’s flash tier sufficient for your latency and quality needs? And if you are running massive token volumes, does Qwen’s ultra‑low output price outweigh any potential differences in model performance? Most startups and SaaS products that can tolerate a slight dip in model fidelity will gravitate toward Qwen‑3.7‑flash because the cost differential is stark. Enterprises that have already invested in Moonshot’s ecosystem or need the Kimi‑specific code‑generation variants will stay with Moonshot despite the higher bill. DeepSeek sits comfortably in the middle, offering a respectable balance of price and capability for teams that are price‑conscious but not ready to switch ecosystems.
Pricing verified on 2026-09-12.