Compare LLM API prices and find the best-value AI model
Compare input/output pricing, context and quality across 82 models — GPT, Claude, Gemini, DeepSeek and more — to pick the right API fast.
Featured Models
OpenAI's newest flagship and the catalog's top score (gated access)
Anthropic's newest flagship — top coding & long-task quality
Huge context, fast and cheap multimodal model
New-gen model blending real-time knowledge & reasoning
Open-source value king with rock-bottom cost
Quick price comparison
Sorted by value by default — see who's most worth it at a glance.
| Provider | Model Name | Try | ||||||
|---|---|---|---|---|---|---|---|---|
Phi-4-mini ChatbotAffordableFast | 68 | 250tok/sFast | $0.02 | $0.04 | 1700.0 | 128K | Try | |
qwen3.7-flash ChatbotAffordableFastLong Context | 82 | 150tok/sFast | $0.03 | $0.13 | 630.8 | 1000K | Try | |
qwen3.5 CodingReasoningLong Context | 91 | 85tok/sFast | $0.10 | $0.15 | 606.7 | 1000K | Try | |
Yi-Lightning ChatbotAffordableFast | 78 | 200tok/sFast | $0.14 | $0.14 | 557.1 | 16K | Try | |
mercury-2.5-preview FastAffordableChatbotCodingLong Context | 82 | 1107tok/sFast | $0.04 | $0.15 | 546.7 | 262K | ||
Phi-4 CodingChatbotAffordableFast | 76 | 180tok/sFast | $0.07 | $0.14 | 542.9 | 16K | Try | |
Baichuan4-Air ChatbotAffordableFast | 70 | 180tok/sFast | $0.14 | $0.14 | 514.7 | 32K | Try | |
Hunyuan HY3 ReasoningChatbotAffordable | 87 | 110tok/sFast | $0.06 | $0.21 | 414.3 | 256K | Try | |
qwen-turbo ChatbotAffordableFast | 74 | 240tok/sFast | $0.05 | $0.20 | 370.0 | 1000K | Try | |
Ling-3.0-flash ChatbotAffordableFastLong Context | 80 | 180tok/sFast | $0.07 | $0.22 | 363.6 | 262K | ||
MAI-DS-R1 ReasoningCoding | 90 | 60tok/sMedium | $0.14 | $0.28 | 321.4 | 128K | Try | |
granite-4.2-8b CodingChatbotAffordableFast | 77 | 130tok/sFast | $0.06 | $0.25 | 308.0 | 131K | ||
Hunyuan TurboS ChatbotAffordableFast | 79 | 160tok/sFast | $0.11 | $0.28 | 282.1 | 256K | Try | |
Step 3.5 Flash ChatbotAffordableFast | 79 | 190tok/sFast | $0.10 | $0.30 | 263.3 | 262K | Try | |
Llama-3.3-70B-Instruct ChatbotCoding | 78 | 100tok/sFast | $0.10 | $0.32 | 243.8 | 128K | Try | |
qwen3.8-flash ChatbotAffordableFastLong Context | 90 | 87tok/sFast | $0.15 | $0.47 | 191.5 | 1000K | Try | |
GLM-5.3-Flash ChatbotAffordableFastLong Context | 89 | 114tok/sFast | $0.15 | $0.50 | 178.0 | 1310K | Try | |
qwen3.5-flash ChatbotAffordableFastLong Context | 82 | 200tok/sFast | $0.15 | $0.47 | 174.5 | 1000K | Try | |
deepseek-v4-flash ChatbotAffordableFastLong Context | 88 | 180tok/sFast | $0.15 | $0.60 | 146.7 | 1000K | Try | |
gpt-4o-mini ChatbotAffordableFast | 80 | 160tok/sFast | $0.15 | $0.60 | 133.3 | 128K | Try | |
qwen-plus ChatbotCodingAffordable | 84 | 110tok/sFast | $0.26 | $0.78 | 107.7 | 131K | Try | |
deepseek-reasoner-v4 ReasoningCoding | 92 | 30tok/sSlow | $0.43 | $0.87 | 105.7 | 131K | Try | |
deepseek-chat-v4 ChatbotAffordableFast | 88 | 140tok/sFast | $0.32 | $0.89 | 98.9 | 131K | Try | |
llama-3.3-70b-versatile ChatbotAffordableFast | 78 | 300tok/sFast | $0.59 | $0.79 | 98.7 | 128K | Try | |
codestral-latest Coding | 80 | 120tok/sFast | $0.30 | $0.90 | 88.9 | 256K | Try | |
meta-llama/Llama-3.3-70B-Instruct-Turbo ChatbotAffordable | 78 | 110tok/sFast | $0.88 | $0.88 | 88.6 | 131K | Try | |
ERNIE X1 ReasoningAffordable | 86 | 45tok/sSlow | $0.28 | $1.10 | 78.2 | 128K | Try | |
gpt-5.6-luna ChatbotFastAffordable | 90 | 150tok/sFast | $0.20 | $1.20 | 75.0 | 1049K | Try | |
MiniMax-M3 CodingReasoningLong ContextAffordable | 89 | 95tok/sFast | $0.30 | $1.20 | 74.2 | 1000K | Try | |
Step 3.7 Flash ChatbotCodingMultimodal | 85 | 140tok/sFast | $0.20 | $1.15 | 73.9 | 262K | Try | |
qwen3.5-plus ChatbotCodingLong Context | 89 | 100tok/sFast | $0.32 | $1.28 | 69.5 | 1000K | Try | |
Qwen/Qwen2.5-72B-Instruct-Turbo ChatbotCodingAffordable | 79 | 100tok/sFast | $1.20 | $1.20 | 65.8 | 131K | Try | |
gpt-5.4-nano ChatbotAffordableFast | 75 | 200tok/sFast | $0.20 | $1.25 | 60.0 | 1049K | Try | |
gemini-3.1-flash-lite ChatbotAffordableFastMultimodal | 80 | 210tok/sFast | $0.25 | $1.50 | 53.3 | 1049K | Try | |
deepseek-v4-pro ReasoningCodingLong Context | 93 | 94tok/sMedium | $0.66 | $1.98 | 47.0 | 1000K | Try | |
doubao-seed-2.1-turbo ChatbotCodingAffordable | 85 | 130tok/sFast | $0.41 | $2.07 | 41.1 | 256K | Try | |
GLM-4.7 ChatbotCodingAffordable | 87 | 130tok/sFast | $0.60 | $2.20 | 39.5 | 205K | Try | |
ERNIE 4.5 ChatbotMultimodal | 84 | 90tok/sFast | $0.55 | $2.20 | 38.2 | 128K | Try | |
grok-4.20 ReasoningChatbotLong Context | 90 | 179tok/sFast | $1.25 | $2.50 | 36.0 | 1000K | Try | |
grok-4.3 ChatbotMultimodalLong Context | 89 | 100tok/sFast | $1.25 | $2.50 | 35.6 | 1000K | Try | |
ERNIE 5.1 ChatbotReasoning | 90 | 80tok/sFast | $0.59 | $2.65 | 34.0 | 128K | Try | |
gemini-2.5-flash ChatbotAffordableFastMultimodal | 84 | 200tok/sFast | $0.30 | $2.50 | 33.6 | 1049K | Try | |
gemini-3.8-flash ReasoningCodingMultimodalLong Context | 92 | 302tok/sFast | $0.75 | $3.75 | 24.5 | 1049K | Try | |
gemini-3.7-flash ReasoningCodingMultimodalLong Context | 90 | 288tok/sFast | $0.75 | $3.75 | 24.0 | 1049K | Try | |
gemini-3.6-flash ChatbotMultimodalLong Context | 85 | 193tok/sFast | $0.75 | $3.75 | 22.7 | 1049K | Try | |
kimi-k2.6 CodingReasoningLong Context | 90 | 85tok/sFast | $0.95 | $4.00 | 22.5 | 262K | Try | |
doubao-seed-2.1-pro ReasoningCodingMultimodal | 91 | 65tok/sMedium | $0.83 | $4.14 | 22.0 | 256K | Try | |
GLM-5.3 CodingReasoningLong Context | 94 | 73tok/sMedium | $1.40 | $4.40 | 21.4 | 1310K | Try | |
GLM-5.2 CodingReasoningLong Context | 93 | 100tok/sFast | $1.40 | $4.40 | 21.1 | 1000K | Try | |
o4-mini ReasoningCodingAffordable | 88 | 80tok/sMedium | $1.10 | $4.40 | 20.0 | 200K | Try | |
gpt-5.4-mini ChatbotCodingAffordable | 85 | 150tok/sFast | $0.75 | $4.50 | 18.9 | 1049K | Try | |
claude-haiku-4-5 ChatbotAffordableFast | 82 | 180tok/sFast | $1.00 | $5.00 | 16.4 | 200K | Try | |
qwen3.8-max ReasoningCodingLong Context | 94 | 70tok/sMedium | $2.00 | $6.00 | 15.7 | 1000K | Try | |
grok-4.6 ReasoningCodingChatbot | 93 | 59tok/sSlow | $2.00 | $6.00 | 15.5 | 500K | Try | |
grok-4.5 ReasoningCodingChatbot | 91 | 91tok/sFast | $2.00 | $6.00 | 15.2 | 500K | Try | |
mistral-large-3 ChatbotCoding | 82 | 80tok/sMedium | $2.00 | $6.00 | 13.7 | 128K | Try | |
o3 ReasoningCoding | 95 | 40tok/sSlow | $2.00 | $8.00 | 11.9 | 200K | Try | |
mistral-medium-3.5 ChatbotCoding | 86 | 90tok/sFast | $1.50 | $7.50 | 11.5 | 128K | Try | |
gemini-3.5-flash ChatbotFastMultimodal | 86 | 190tok/sFast | $1.50 | $9.00 | 9.6 | 1049K | Try | |
gemini-2.5-pro CodingReasoningLong ContextMultimodal | 93 | 65tok/sMedium | $1.25 | $10.00 | 9.3 | 1049K | Try | |
claude-sonnet-5 CodingReasoningLong Context | 92 | 95tok/sFast | $2.00 | $10.00 | 9.2 | 1049K | Try | |
gpt-4o ChatbotMultimodal | 88 | 100tok/sFast | $2.50 | $10.00 | 8.8 | 128K | Try | |
command-a ChatbotLong Context | 83 | 70tok/sMedium | $2.50 | $10.00 | 8.3 | 256K | Try | |
gpt-5.6-terra ChatbotCodingMultimodal | 95 | 110tok/sFast | $2.00 | $12.00 | 7.9 | 1049K | Try | |
gemini-3.1-pro CodingReasoningLong ContextMultimodal | 91 | 70tok/sMedium | $2.00 | $12.00 | 7.6 | 1049K | Try | |
command-r ChatbotAffordable | 75 | 120tok/sFast | $2.50 | $10.00 | 7.5 | 128K | Try | |
kimi-k3 ReasoningCodingLong Context | 95 | 60tok/sMedium | $3.00 | $15.00 | 6.3 | 1000K | Try | |
gpt-5.4 ChatbotCodingMultimodal | 92 | 95tok/sFast | $2.50 | $15.00 | 6.1 | 1049K | Try | |
claude-sonnet-4-6 CodingChatbot | 90 | 90tok/sFast | $3.00 | $15.00 | 6.0 | 1049K | Try | |
Baichuan4 ChatbotReasoning | 80 | 60tok/sMedium | $13.89 | $13.89 | 5.8 | 32K | Try | |
gpt-5.6-sol ReasoningCodingMultimodal | 98 | 85tok/sMedium | $4.00 | $20.00 | 4.9 | 1049K | Try | |
claude-opus-5 CodingReasoningLong Context | 97 | 58tok/sMedium | $5.00 | $25.00 | 3.9 | 1049K | Try | |
claude-opus-4-8 CodingReasoningLong Context | 94 | 50tok/sMedium | $5.00 | $25.00 | 3.8 | 1049K | Try | |
claude-opus-4-6 CodingReasoningLong Context | 93 | 45tok/sSlow | $5.00 | $25.00 | 3.7 | 1049K | Try | |
gpt-5.5 ChatbotReasoningMultimodal | 97 | 85tok/sMedium | $5.00 | $30.00 | 3.2 | 1049K | Try | |
gpt-6-astra ReasoningCodingLong ContextMultimodal | 99 | 58tok/sMedium | $10.00 | $50.00 | 2.0 | 1050K | Try | |
claude-fable-5.1 CodingReasoningLong Context | 97 | 66tok/sSlow | $10.00 | $50.00 | 1.9 | 1000K | Try | |
claude-fable-5 CodingReasoning | 96 | 60tok/sMedium | $10.00 | $50.00 | 1.9 | 1049K | Try | |
claude-mythos-5 ChatbotReasoning | 96 | 55tok/sMedium | $10.00 | $50.00 | 1.9 | 1049K | Try | |
gpt-6-luna | N/A | N/A | $0.10 | $0.50 | N/A | 1050K | Try | |
gpt-6-sol | N/A | N/A | $2.00 | $10.00 | N/A | 1050K | Try | |
claude-opus-5.5 | N/A | N/A | $4.00 | $20.00 | N/A | 1000K | Try |
Choose a model by your scenario
Not sure where to start? Jump in from the three most common needs.
Coding: which model writes the best code?
Compare code understanding, tool use and long-task reliability to find the strongest coding model.
See coding modelsRAG & long context: compare large-context models
Weigh context length, input cost and retrieval strategy to find the right long-context model.
See long-context modelsHigh throughput, low cost: choosing a production API
Work backward from monthly volume, latency and token rates to the most affordable pick.
See budget modelsRecently added and updated models
Prices and models are refreshed continuously by our daily data pipeline.
gpt-6-astra
In 10.00 · Out 50.00 / 1M tokens
mercury-2.5-preview
In 0.04 · Out 0.15 / 1M tokens
gemini-3.8-flash
In 0.75 · Out 3.75 / 1M tokens
granite-4.2-8b
In 0.06 · Out 0.25 / 1M tokens
GLM-5.3-Flash
In 0.15 · Out 0.50 / 1M tokens
deepseek-v4-flash
In 0.15 · Out 0.60 / 1M tokens
claude-fable-5.1
In 10.00 · Out 50.00 / 1M tokens
grok-4.6
In 2.00 · Out 6.00 / 1M tokens
qwen3.7-flash
In 0.03 · Out 0.13 / 1M tokens
GLM-5.3
In 1.40 · Out 4.40 / 1M tokens
claude-opus-5
In 5.00 · Out 25.00 / 1M tokens
gpt-5.6-sol
In 4.00 · Out 20.00 / 1M tokens
grok-4.5
In 2.00 · Out 6.00 / 1M tokens
claude-sonnet-5
In 2.00 · Out 10.00 / 1M tokens
Questions you might have about LLM API pricing
What's the cheapest LLM API in 2026?+
By per-million-token pricing, the floor in this catalog is Microsoft's Phi-4-mini at $0.02/$0.04 — though it is a small model with only 128K of context. For the cheapest option that still gives you a million-token window, Alibaba's qwen3.7-flash is $0.03/$0.13, with Inception's mercury-2.5-preview close behind at $0.04/$0.15 over 262K. Raise the quality bar to 90 and the cheapest is qwen3.5 at $0.10/$0.15. Output rates down here run more than a hundred times below the flagship tier, so sort the table above by price and cross-check the value score. Note that launch promotions expire: every row on this site carries its price source and verification date, so budget long-term against the list rate.
GPT-5.5 or Claude Sonnet 5 — which is better?+
GPT-5.5 scores higher on multimodal and overall reasoning, ideal for flagship tasks mixing text, image and audio. Claude Sonnet 5 excels at coding and long-task reliability at a friendlier price. Favor Claude for code-heavy work and GPT-5.5 for an all-round flagship, then compare quality and cost on your actual tasks.
How is LLM API pricing calculated? What are input/output tokens?+
LLM APIs bill separately per million tokens for input (what you send) and output (what the model generates), with output usually several times more expensive. Monthly cost ≈ (avg input tokens × rate + avg output tokens × rate) × monthly requests. Use our cost calculator to estimate it directly.
Which LLM is best for RAG applications?+
RAG depends on both context window size and input token price, since retrieved documents consume a lot of input. Gemini's huge context plus low input pricing suits RAG well; for smaller corpora, a cheaper mid-tier model with good retrieval and caching keeps costs down.
How do I choose the right AI model?+
Define your quality bar first, then compare input/output pricing, latency and speed, context length and modalities. Start from a scenario (coding, RAG, production), narrow down with the comparison table and recommendation wizard, and validate your monthly budget with the cost calculator.
Explore more tools and content
From the model encyclopedia to rankings and guides, find everything you need.
Models
Explore every model by capability, price and scenario
Rankings
Value, cheapest and top-quality rankings
Calculator
Estimate your monthly API cost
Recommend
Answer three questions to find your best fit
Providers
Models and pricing across every provider
Blog
Guides on comparing, choosing and saving on AI models