LLM Insights中文
Cohere logo

Cohere API Pricing & Model Comparison

We track 2 Cohere models, with input pricing from $2.50 to $2.50 per million tokens. The line-up's blended rate runs 60% above the catalog average. Below: price, context, speed and value compared across the range.

Models tracked2
Lowest input price$2.50 / 1M
Top quality score83
Largest context256K

For cost, Cohere's cheapest option is command-a ($2.50 input / $10.00 output); the quality ceiling is command-a at a quality score of 83. On a 50,000-turn-per-month support workload that is $500 versus $500 — the spread inside a single provider is often wider than the spread between providers.

On value (quality score ÷ output price) the pick of the Cohere range is command-a. The largest context window is 256K, offered only by command-a, and the fastest output is command-r at 120 tok/s. Modalities across the line-up: Text.

Models and pricing

Sorted by blended rate, cheapest first (3:1 input-to-output mix). Follow a model name for full specs and monthly cost estimates.

ModelInput /1MOutput /1MBlended /1MQualityContextSpeedValueUse cases
command-a$2.50$10.00$4.3883256K70 tok/s8.3
ChatbotLong Context
command-r$2.50$10.00$4.3875128K120 tok/s7.5
ChatbotAffordable

Frequently asked questions

Which Cohere model is cheapest?

command-a, at $2.50 input and $10.00 output per million tokens, with a quality score of 83.

What is Cohere's flagship model?

By quality score it is command-a (83), priced at $2.50 input / $10.00 output with a 256K context window.

Does Cohere pricing change?

Yes. llmprice.app runs a daily collection job and refreshes these figures, but there can be a lag after a provider re-prices — confirm against Cohere's official pricing page before committing.

Browse all providers

Prices and specs are collected daily by llmprice.app. Confirm against the official pricing page before production use.