Cohere API Pricing & Model Comparison
We track 2 Cohere models, with input pricing from $2.50 to $2.50 per million tokens. The line-up's blended rate runs 60% above the catalog average. Below: price, context, speed and value compared across the range.
For cost, Cohere's cheapest option is command-a ($2.50 input / $10.00 output); the quality ceiling is command-a at a quality score of 83. On a 50,000-turn-per-month support workload that is $500 versus $500 — the spread inside a single provider is often wider than the spread between providers.
On value (quality score ÷ output price) the pick of the Cohere range is command-a. The largest context window is 256K, offered only by command-a, and the fastest output is command-r at 120 tok/s. Modalities across the line-up: Text.
Models and pricing
Sorted by blended rate, cheapest first (3:1 input-to-output mix). Follow a model name for full specs and monthly cost estimates.
Frequently asked questions
Which Cohere model is cheapest?
command-a, at $2.50 input and $10.00 output per million tokens, with a quality score of 83.
What is Cohere's flagship model?
By quality score it is command-a (83), priced at $2.50 input / $10.00 output with a 256K context window.
Does Cohere pricing change?
Yes. llmprice.app runs a daily collection job and refreshes these figures, but there can be a lag after a provider re-prices — confirm against Cohere's official pricing page before committing.
Prices and specs are collected daily by llmprice.app. Confirm against the official pricing page before production use.