LLM Insights中文
IBM logo

IBM API Pricing & Model Comparison

We track 1 IBM models, with input pricing from $0.06 to $0.06 per million tokens. The line-up's blended rate runs 97% below the catalog average. Below: price, context, speed and value compared across the range.

Models tracked1
Lowest input price$0.06 / 1M
Top quality score77
Largest context131K

For cost, IBM's cheapest option is granite-4.2-8b ($0.06 input / $0.25 output); the quality ceiling is granite-4.2-8b at a quality score of 77. On a 50,000-turn-per-month support workload that is $12 versus $12 — the spread inside a single provider is often wider than the spread between providers.

On value (quality score ÷ output price) the pick of the IBM range is granite-4.2-8b. The largest context window is 131K, offered only by granite-4.2-8b, and the fastest output is granite-4.2-8b at 130 tok/s. Modalities across the line-up: Text.

Models and pricing

Sorted by blended rate, cheapest first (3:1 input-to-output mix). Follow a model name for full specs and monthly cost estimates.

ModelInput /1MOutput /1MBlended /1MQualityContextSpeedValueUse cases
granite-4.2-8b$0.06$0.25$0.1177131K130 tok/s308.0
CodingChatbotAffordableFast

Frequently asked questions

Which IBM model is cheapest?

granite-4.2-8b, at $0.06 input and $0.25 output per million tokens, with a quality score of 77.

What is IBM's flagship model?

By quality score it is granite-4.2-8b (77), priced at $0.06 input / $0.25 output with a 131K context window.

Does IBM pricing change?

Yes. llmprice.app runs a daily collection job and refreshes these figures, but there can be a lag after a provider re-prices — confirm against IBM's official pricing page before committing.

Browse all providers

Prices and specs are collected daily by llmprice.app. Confirm against the official pricing page before production use.