Cohere’s command‑r7b‑12‑2024 undercuts OpenAI’s gpt‑4o‑mini on input cost while OpenAI wins on output, a split that flips the economics depending on how verbose your answers need to be.
How the token meter reads for RAG‑focused models
| Vendor | Model | Input | Output |
|---|---|---|---|
| Cohere | command‑r7b‑12‑2024 | $0.0375 per 1M tokens | $0.15 per 1M tokens |
| Cohere | command‑r‑08‑2024 | $0.15 per 1M tokens | $0.6 per 1M tokens |
| Cohere | command‑a | $2.5 per 1M tokens | $10 per 1M tokens |
| Cohere | command‑r‑plus‑08‑2024 | $2.5 per 1M tokens | $10 per 1M tokens |
| OpenAI | gpt‑4o‑mini | $0.15 per 1M tokens | $0.6 per 1M tokens |
| OpenAI | gpt‑4o‑mini‑2024‑07‑18 | $0.15 per 1M tokens | $0.6 per 1M tokens |
The table strips away marketing fluff and shows the raw per‑token rates that land on your invoice. Cohere’s cheapest RAG candidate, command‑r7b‑12‑2024, charges a fraction of a cent for input but still costs less than a quarter of a cent for output. OpenAI’s gpt‑4o‑mini charges a higher input rate but matches Cohere’s output fee at $0.6 per 1M tokens. The higher‑tier Cohere models—command‑a and command‑r‑plus‑08‑2024—jump to $2.5 per 1M input tokens and $10 per 1M output tokens, a steep premium that only makes sense for workloads demanding the most advanced reasoning or strict safety guarantees.
Choosing the model that fits your token profile
If your RAG pipeline primarily retrieves short passages and returns concise answers, the lower input cost of Cohere’s command‑r7b‑12‑2024 can shave dollars off each query. For use cases that generate long‑form content—legal briefs, technical reports, or detailed summaries—the output fee becomes the dominant factor, and OpenAI’s gpt‑4o‑mini’s $0.6 per 1M output tokens can keep the bill tighter than Cohere’s $0.15 output rate. Caching also tilts the balance: OpenAI lists cached‑input discounts for other models (for example, $0.005 cached input for gpt‑5‑nano), but Cohere does not publish a cached‑input rate, meaning any reuse of retrieved chunks will be billed at the full input price on the Cohere side.
When the shape of your token consumption leans toward heavy output, Cohere’s lower output price makes it the pragmatic pick. When the workload is input‑heavy or you rely on cached prompts, OpenAI’s broader pricing ecosystem and cached‑input options may deliver a leaner bill. Both vendors claim comparable quality for classification and RAG tasks, so the decision hinges on whether you value input economy, output economy, or the flexibility of cached pricing.
Pricing verified on 2026-09-12.