Compare LLM API prices and find the best-value AI model
Compare input/output pricing, context and quality across 30+ models — GPT, Claude, Gemini, DeepSeek and more — to pick the right API fast.
Featured Models
All-round flagship with top reasoning & multimodal skills
Anthropic's newest flagship — top coding & long-task quality
Huge context, fast and cheap multimodal model
New-gen model blending real-time knowledge & reasoning
Open-source value king with rock-bottom cost
Quick price comparison
Sorted by value by default — see who's most worth it at a glance.
| Provider | Model Name | Try | ||||||
|---|---|---|---|---|---|---|---|---|
Phi-4-mini ChatbotAffordableFast | 68 | 250tok/sFast | $0.02 | $0.04 | 1700.0 | 128K | Try | |
Yi-Lightning ChatbotAffordableFast | 78 | 200tok/sFast | $0.14 | $0.14 | 557.1 | 16K | Try | |
Phi-4 CodingChatbotAffordableFast | 76 | 180tok/sFast | $0.07 | $0.14 | 542.9 | 16K | Try | |
Baichuan4-Air ChatbotAffordableFast | 70 | 180tok/sFast | $0.14 | $0.14 | 514.7 | 32K | Try | |
MAI-DS-R1 ReasoningCoding | 90 | 60tok/sMedium | $0.14 | $0.28 | 321.4 | 128K | Try | |
Llama-3.3-70B-Instruct ChatbotCoding | 78 | 100tok/sFast | $0.10 | $0.32 | 243.8 | 128K | Try | |
gpt-4o-mini ChatbotAffordableFast | 80 | 160tok/sFast | $0.15 | $0.60 | 133.3 | 128K | Try | |
qwen-plus ChatbotCodingAffordable | 84 | 110tok/sFast | $0.26 | $0.78 | 107.7 | 131K | Try | |
deepseek-reasoner-v4 ReasoningCoding | 92 | 30tok/sSlow | $0.43 | $0.87 | 105.7 | 131K | Try | |
llama-3.3-70b-versatile ChatbotAffordableFast | 78 | 300tok/sFast | $0.59 | $0.79 | 98.7 | 128K | Try | |
meta-llama/Llama-3.3-70B-Instruct-Turbo ChatbotAffordable | 78 | 110tok/sFast | $0.88 | $0.88 | 88.6 | 131K | Try | |
deepseek-chat-v4 ChatbotAffordableFast | 88 | 140tok/sFast | $0.26 | $1.03 | 85.5 | 131K | Try | |
ERNIE X1 ReasoningAffordable | 86 | 45tok/sSlow | $0.28 | $1.10 | 78.2 | 128K | Try | |
Qwen/Qwen2.5-72B-Instruct-Turbo ChatbotCodingAffordable | 79 | 100tok/sFast | $1.20 | $1.20 | 65.8 | 131K | Try | |
claude-haiku-4-5 ChatbotAffordableFast | 82 | 180tok/sFast | $1.00 | $1.25 | 65.6 | 200K | Try | |
ERNIE 4.5 ChatbotMultimodal | 84 | 90tok/sFast | $0.55 | $2.20 | 38.2 | 128K | Try | |
ERNIE 5.1 ChatbotReasoning | 90 | 80tok/sFast | $0.59 | $2.65 | 34.0 | 128K | Try | |
o4-mini ReasoningCodingAffordable | 88 | 80tok/sMedium | $1.10 | $4.40 | 20.0 | 200K | Try | |
grok-4.3 ChatbotMultimodal | 89 | 100tok/sFast | $2.00 | $6.00 | 14.8 | 131K | Try | |
mistral-large-3 ChatbotCoding | 82 | 80tok/sMedium | $2.00 | $6.00 | 13.7 | 128K | Try | |
Mistral-Large ChatbotCoding | 82 | 80tok/sMedium | $2.00 | $6.00 | 13.7 | 128K | Try | |
o3 ReasoningCoding | 95 | 40tok/sSlow | $2.00 | $8.00 | 11.9 | 200K | Try | |
mistral-medium-3.5 ChatbotCoding | 86 | 90tok/sFast | $1.50 | $7.50 | 11.5 | 128K | Try | |
gpt-4o ChatbotMultimodal | 88 | 100tok/sFast | $2.50 | $10.00 | 8.8 | 128K | Try | |
command-r ChatbotAffordable | 75 | 120tok/sFast | $2.50 | $10.00 | 7.5 | 128K | Try | |
Baichuan4 ChatbotReasoning | 80 | 60tok/sMedium | $13.89 | $13.89 | 5.8 | 32K | Try |
Choose a model by your scenario
Not sure where to start? Jump in from the three most common needs.
Coding: which model writes the best code?
Compare code understanding, tool use and long-task reliability to find the strongest coding model.
See coding modelsRAG & long context: compare large-context models
Weigh context length, input cost and retrieval strategy to find the right long-context model.
See long-context modelsHigh throughput, low cost: choosing a production API
Work backward from monthly volume, latency and token rates to the most affordable pick.
See budget modelsRecently added and updated models
Prices and models are refreshed continuously by our daily data pipeline.
claude-opus-5
In 5.00 · Out 6.25 / 1M tokens
gpt-5.6-sol
In 5.00 · Out 30.00 / 1M tokens
gpt-5.6-terra
In 1.00 · Out 6.00 / 1M tokens
grok-4.5
In 2.00 · Out 6.00 / 1M tokens
claude-sonnet-5
In 3.00 · Out 3.75 / 1M tokens
gemini-3.1-flash-lite
In 0.25 · Out 1.50 / 1M tokens
GLM-5.2
In 1.40 · Out 4.40 / 1M tokens
Questions you might have about LLM API pricing
What's the cheapest LLM API in 2026?+
By per-million-token pricing, open-source models like DeepSeek, plus Gemini Flash Lite and GLM, are usually the cheapest — output prices can be 10× lower than flagship models. The best choice depends on your quality bar, so sort the table above by price and cross-check the value score.
GPT-5.5 or Claude Sonnet 5 — which is better?+
GPT-5.5 scores higher on multimodal and overall reasoning, ideal for flagship tasks mixing text, image and audio. Claude Sonnet 5 excels at coding and long-task reliability at a friendlier price. Favor Claude for code-heavy work and GPT-5.5 for an all-round flagship, then compare quality and cost on your actual tasks.
How is LLM API pricing calculated? What are input/output tokens?+
LLM APIs bill separately per million tokens for input (what you send) and output (what the model generates), with output usually several times more expensive. Monthly cost ≈ (avg input tokens × rate + avg output tokens × rate) × monthly requests. Use our cost calculator to estimate it directly.
Which LLM is best for RAG applications?+
RAG depends on both context window size and input token price, since retrieved documents consume a lot of input. Gemini's huge context plus low input pricing suits RAG well; for smaller corpora, a cheaper mid-tier model with good retrieval and caching keeps costs down.
How do I choose the right AI model?+
Define your quality bar first, then compare input/output pricing, latency and speed, context length and modalities. Start from a scenario (coding, RAG, production), narrow down with the comparison table and recommendation wizard, and validate your monthly budget with the cost calculator.
Explore more tools and content
From the model encyclopedia to rankings and guides, find everything you need.
Models
Explore every model by capability, price and scenario
Rankings
Value, cheapest and top-quality rankings
Calculator
Estimate your monthly API cost
Recommend
Answer three questions to find your best fit
Providers
Models and pricing across every provider
Blog
Guides on comparing, choosing and saving on AI models