IBM API Pricing & Model Comparison
We track 1 IBM models, with input pricing from $0.06 to $0.06 per million tokens. The line-up's blended rate runs 97% below the catalog average. Below: price, context, speed and value compared across the range.
For cost, IBM's cheapest option is granite-4.2-8b ($0.06 input / $0.25 output); the quality ceiling is granite-4.2-8b at a quality score of 77. On a 50,000-turn-per-month support workload that is $12 versus $12 — the spread inside a single provider is often wider than the spread between providers.
On value (quality score ÷ output price) the pick of the IBM range is granite-4.2-8b. The largest context window is 131K, offered only by granite-4.2-8b, and the fastest output is granite-4.2-8b at 130 tok/s. Modalities across the line-up: Text.
Models and pricing
Sorted by blended rate, cheapest first (3:1 input-to-output mix). Follow a model name for full specs and monthly cost estimates.
| Model | Input /1M | Output /1M | Blended /1M | Quality | Context | Speed | Value | Use cases |
|---|---|---|---|---|---|---|---|---|
| granite-4.2-8b | $0.06 | $0.25 | $0.11 | 77 | 131K | 130 tok/s | 308.0 | CodingChatbotAffordableFast |
Frequently asked questions
Which IBM model is cheapest?
granite-4.2-8b, at $0.06 input and $0.25 output per million tokens, with a quality score of 77.
What is IBM's flagship model?
By quality score it is granite-4.2-8b (77), priced at $0.06 input / $0.25 output with a 131K context window.
Does IBM pricing change?
Yes. llmprice.app runs a daily collection job and refreshes these figures, but there can be a lag after a provider re-prices — confirm against IBM's official pricing page before committing.
Prices and specs are collected daily by llmprice.app. Confirm against the official pricing page before production use.