LLM Insights中文

Compare LLM API prices and find the best-value AI model

Compare input/output pricing, context and quality across 82 models — GPT, Claude, Gemini, DeepSeek and more — to pick the right API fast.

Featured Models

OpenAI logo
gpt-6-astra
OpenAI
99

OpenAI's newest flagship and the catalog's top score (gated access)

Input Price$10.00/1M tokens
Output Price$50.00/1M tokens
Context Window1050K
Output Speed58 tok/sMedium
Text
Vision
Anthropic logo
claude-opus-5
Anthropic
97

Anthropic's newest flagship — top coding & long-task quality

Input Price$5.00/1M tokens
Output Price$25.00/1M tokens
Context Window1049K
Output Speed58 tok/sMedium
Text
Vision
Google logo
gemini-3.5-flash
Google
86

Huge context, fast and cheap multimodal model

Input Price$1.50/1M tokens
Output Price$9.00/1M tokens
Context Window1049K
Output Speed190 tok/sFast
Text
Vision
Audio
xAI logo
grok-4.3
xAI
89

New-gen model blending real-time knowledge & reasoning

Input Price$1.25/1M tokens
Output Price$2.50/1M tokens
Context Window1000K
Output Speed100 tok/sFast
Text
Vision
DeepSeek logo
deepseek-chat-v4
DeepSeek
88

Open-source value king with rock-bottom cost

Input Price$0.32/1M tokens
Output Price$0.89/1M tokens
Context Window131K
Output Speed140 tok/sFast
Text

Quick price comparison

Sorted by value by default — see who's most worth it at a glance.

ProviderModel NameTry
Microsoft logoMicrosoft
Phi-4-mini
ChatbotAffordableFast
68
250tok/sFast
$0.02$0.041700.0128KTry
Alibaba logoAlibaba
qwen3.7-flash
ChatbotAffordableFastLong Context
82
150tok/sFast
$0.03$0.13630.81000KTry
Alibaba logoAlibaba
qwen3.5
CodingReasoningLong Context
91
85tok/sFast
$0.10$0.15606.71000KTry
01.AI logo01.AI
Yi-Lightning
ChatbotAffordableFast
78
200tok/sFast
$0.14$0.14557.116KTry
Inception logoInception
mercury-2.5-preview
FastAffordableChatbotCodingLong Context
82
1107tok/sFast
$0.04$0.15546.7262K
Microsoft logoMicrosoft
Phi-4
CodingChatbotAffordableFast
76
180tok/sFast
$0.07$0.14542.916KTry
Baichuan logoBaichuan
Baichuan4-Air
ChatbotAffordableFast
70
180tok/sFast
$0.14$0.14514.732KTry
Tencent logoTencent
Hunyuan HY3
ReasoningChatbotAffordable
87
110tok/sFast
$0.06$0.21414.3256KTry
Alibaba logoAlibaba
qwen-turbo
ChatbotAffordableFast
74
240tok/sFast
$0.05$0.20370.01000KTry
InclusionAI logoInclusionAI
Ling-3.0-flash
ChatbotAffordableFastLong Context
80
180tok/sFast
$0.07$0.22363.6262K
Microsoft logoMicrosoft
MAI-DS-R1
ReasoningCoding
90
60tok/sMedium
$0.14$0.28321.4128KTry
IBM logoIBM
granite-4.2-8b
CodingChatbotAffordableFast
77
130tok/sFast
$0.06$0.25308.0131K
Tencent logoTencent
Hunyuan TurboS
ChatbotAffordableFast
79
160tok/sFast
$0.11$0.28282.1256KTry
StepFun logoStepFun
Step 3.5 Flash
ChatbotAffordableFast
79
190tok/sFast
$0.10$0.30263.3262KTry
Microsoft logoMicrosoft
Llama-3.3-70B-Instruct
ChatbotCoding
78
100tok/sFast
$0.10$0.32243.8128KTry
Alibaba logoAlibaba
qwen3.8-flash
ChatbotAffordableFastLong Context
90
87tok/sFast
$0.15$0.47191.51000KTry
Zhipu logoZhipu
GLM-5.3-Flash
ChatbotAffordableFastLong Context
89
114tok/sFast
$0.15$0.50178.01310KTry
Alibaba logoAlibaba
qwen3.5-flash
ChatbotAffordableFastLong Context
82
200tok/sFast
$0.15$0.47174.51000KTry
DeepSeek logoDeepSeek
deepseek-v4-flash
ChatbotAffordableFastLong Context
88
180tok/sFast
$0.15$0.60146.71000KTry
OpenAI logoOpenAI
gpt-4o-mini
ChatbotAffordableFast
80
160tok/sFast
$0.15$0.60133.3128KTry
Alibaba logoAlibaba
qwen-plus
ChatbotCodingAffordable
84
110tok/sFast
$0.26$0.78107.7131KTry
DeepSeek logoDeepSeek
deepseek-reasoner-v4
ReasoningCoding
92
30tok/sSlow
$0.43$0.87105.7131KTry
DeepSeek logoDeepSeek
deepseek-chat-v4
ChatbotAffordableFast
88
140tok/sFast
$0.32$0.8998.9131KTry
Groq logoGroq
llama-3.3-70b-versatile
ChatbotAffordableFast
78
300tok/sFast
$0.59$0.7998.7128KTry
Mistral logoMistral
codestral-latest
Coding
80
120tok/sFast
$0.30$0.9088.9256KTry
Together AI logoTogether AI
meta-llama/Llama-3.3-70B-Instruct-Turbo
ChatbotAffordable
78
110tok/sFast
$0.88$0.8888.6131KTry
Baidu logoBaidu
ERNIE X1
ReasoningAffordable
86
45tok/sSlow
$0.28$1.1078.2128KTry
OpenAI logoOpenAI
gpt-5.6-luna
ChatbotFastAffordable
90
150tok/sFast
$0.20$1.2075.01049KTry
MiniMax logoMiniMax
MiniMax-M3
CodingReasoningLong ContextAffordable
89
95tok/sFast
$0.30$1.2074.21000KTry
StepFun logoStepFun
Step 3.7 Flash
ChatbotCodingMultimodal
85
140tok/sFast
$0.20$1.1573.9262KTry
Alibaba logoAlibaba
qwen3.5-plus
ChatbotCodingLong Context
89
100tok/sFast
$0.32$1.2869.51000KTry
Together AI logoTogether AI
Qwen/Qwen2.5-72B-Instruct-Turbo
ChatbotCodingAffordable
79
100tok/sFast
$1.20$1.2065.8131KTry
OpenAI logoOpenAI
gpt-5.4-nano
ChatbotAffordableFast
75
200tok/sFast
$0.20$1.2560.01049KTry
Google logoGoogle
gemini-3.1-flash-lite
ChatbotAffordableFastMultimodal
80
210tok/sFast
$0.25$1.5053.31049KTry
DeepSeek logoDeepSeek
deepseek-v4-pro
ReasoningCodingLong Context
93
94tok/sMedium
$0.66$1.9847.01000KTry
ByteDance logoByteDance
doubao-seed-2.1-turbo
ChatbotCodingAffordable
85
130tok/sFast
$0.41$2.0741.1256KTry
Zhipu logoZhipu
GLM-4.7
ChatbotCodingAffordable
87
130tok/sFast
$0.60$2.2039.5205KTry
Baidu logoBaidu
ERNIE 4.5
ChatbotMultimodal
84
90tok/sFast
$0.55$2.2038.2128KTry
xAI logoxAI
grok-4.20
ReasoningChatbotLong Context
90
179tok/sFast
$1.25$2.5036.01000KTry
xAI logoxAI
grok-4.3
ChatbotMultimodalLong Context
89
100tok/sFast
$1.25$2.5035.61000KTry
Baidu logoBaidu
ERNIE 5.1
ChatbotReasoning
90
80tok/sFast
$0.59$2.6534.0128KTry
Google logoGoogle
gemini-2.5-flash
ChatbotAffordableFastMultimodal
84
200tok/sFast
$0.30$2.5033.61049KTry
Google logoGoogle
gemini-3.8-flash
ReasoningCodingMultimodalLong Context
92
302tok/sFast
$0.75$3.7524.51049KTry
Google logoGoogle
gemini-3.7-flash
ReasoningCodingMultimodalLong Context
90
288tok/sFast
$0.75$3.7524.01049KTry
Google logoGoogle
gemini-3.6-flash
ChatbotMultimodalLong Context
85
193tok/sFast
$0.75$3.7522.71049KTry
Moonshot AI logoMoonshot AI
kimi-k2.6
CodingReasoningLong Context
90
85tok/sFast
$0.95$4.0022.5262KTry
ByteDance logoByteDance
doubao-seed-2.1-pro
ReasoningCodingMultimodal
91
65tok/sMedium
$0.83$4.1422.0256KTry
Zhipu logoZhipu
GLM-5.3
CodingReasoningLong Context
94
73tok/sMedium
$1.40$4.4021.41310KTry
Zhipu logoZhipu
GLM-5.2
CodingReasoningLong Context
93
100tok/sFast
$1.40$4.4021.11000KTry
OpenAI logoOpenAI
o4-mini
ReasoningCodingAffordable
88
80tok/sMedium
$1.10$4.4020.0200KTry
OpenAI logoOpenAI
gpt-5.4-mini
ChatbotCodingAffordable
85
150tok/sFast
$0.75$4.5018.91049KTry
Anthropic logoAnthropic
claude-haiku-4-5
ChatbotAffordableFast
82
180tok/sFast
$1.00$5.0016.4200KTry
Alibaba logoAlibaba
qwen3.8-max
ReasoningCodingLong Context
94
70tok/sMedium
$2.00$6.0015.71000KTry
xAI logoxAI
grok-4.6
ReasoningCodingChatbot
93
59tok/sSlow
$2.00$6.0015.5500KTry
xAI logoxAI
grok-4.5
ReasoningCodingChatbot
91
91tok/sFast
$2.00$6.0015.2500KTry
Mistral logoMistral
mistral-large-3
ChatbotCoding
82
80tok/sMedium
$2.00$6.0013.7128KTry
OpenAI logoOpenAI
o3
ReasoningCoding
95
40tok/sSlow
$2.00$8.0011.9200KTry
Mistral logoMistral
mistral-medium-3.5
ChatbotCoding
86
90tok/sFast
$1.50$7.5011.5128KTry
Google logoGoogle
gemini-3.5-flash
ChatbotFastMultimodal
86
190tok/sFast
$1.50$9.009.61049KTry
Google logoGoogle
gemini-2.5-pro
CodingReasoningLong ContextMultimodal
93
65tok/sMedium
$1.25$10.009.31049KTry
Anthropic logoAnthropic
claude-sonnet-5
CodingReasoningLong Context
92
95tok/sFast
$2.00$10.009.21049KTry
OpenAI logoOpenAI
gpt-4o
ChatbotMultimodal
88
100tok/sFast
$2.50$10.008.8128KTry
Cohere logoCohere
command-a
ChatbotLong Context
83
70tok/sMedium
$2.50$10.008.3256KTry
OpenAI logoOpenAI
gpt-5.6-terra
ChatbotCodingMultimodal
95
110tok/sFast
$2.00$12.007.91049KTry
Google logoGoogle
gemini-3.1-pro
CodingReasoningLong ContextMultimodal
91
70tok/sMedium
$2.00$12.007.61049KTry
Cohere logoCohere
command-r
ChatbotAffordable
75
120tok/sFast
$2.50$10.007.5128KTry
Moonshot AI logoMoonshot AI
kimi-k3
ReasoningCodingLong Context
95
60tok/sMedium
$3.00$15.006.31000KTry
OpenAI logoOpenAI
gpt-5.4
ChatbotCodingMultimodal
92
95tok/sFast
$2.50$15.006.11049KTry
Anthropic logoAnthropic
claude-sonnet-4-6
CodingChatbot
90
90tok/sFast
$3.00$15.006.01049KTry
Baichuan logoBaichuan
Baichuan4
ChatbotReasoning
80
60tok/sMedium
$13.89$13.895.832KTry
OpenAI logoOpenAI
gpt-5.6-sol
ReasoningCodingMultimodal
98
85tok/sMedium
$4.00$20.004.91049KTry
Anthropic logoAnthropic
claude-opus-5
CodingReasoningLong Context
97
58tok/sMedium
$5.00$25.003.91049KTry
Anthropic logoAnthropic
claude-opus-4-8
CodingReasoningLong Context
94
50tok/sMedium
$5.00$25.003.81049KTry
Anthropic logoAnthropic
claude-opus-4-6
CodingReasoningLong Context
93
45tok/sSlow
$5.00$25.003.71049KTry
OpenAI logoOpenAI
gpt-5.5
ChatbotReasoningMultimodal
97
85tok/sMedium
$5.00$30.003.21049KTry
OpenAI logoOpenAI
gpt-6-astra
ReasoningCodingLong ContextMultimodal
99
58tok/sMedium
$10.00$50.002.01050KTry
Anthropic logoAnthropic
claude-fable-5.1
CodingReasoningLong Context
97
66tok/sSlow
$10.00$50.001.91000KTry
Anthropic logoAnthropic
claude-fable-5
CodingReasoning
96
60tok/sMedium
$10.00$50.001.91049KTry
Anthropic logoAnthropic
claude-mythos-5
ChatbotReasoning
96
55tok/sMedium
$10.00$50.001.91049KTry
OpenAI logoOpenAI
gpt-6-luna
N/AN/A$0.10$0.50N/A1050KTry
OpenAI logoOpenAI
gpt-6-sol
N/AN/A$2.00$10.00N/A1050KTry
Anthropic logoAnthropic
claude-opus-5.5
N/AN/A$4.00$20.00N/A1000KTry
Guided picks

Choose a model by your scenario

Not sure where to start? Jump in from the three most common needs.

What's new

Recently added and updated models

Prices and models are refreshed continuously by our daily data pipeline.

View all models
FAQ

Questions you might have about LLM API pricing

What's the cheapest LLM API in 2026?+

By per-million-token pricing, the floor in this catalog is Microsoft's Phi-4-mini at $0.02/$0.04 — though it is a small model with only 128K of context. For the cheapest option that still gives you a million-token window, Alibaba's qwen3.7-flash is $0.03/$0.13, with Inception's mercury-2.5-preview close behind at $0.04/$0.15 over 262K. Raise the quality bar to 90 and the cheapest is qwen3.5 at $0.10/$0.15. Output rates down here run more than a hundred times below the flagship tier, so sort the table above by price and cross-check the value score. Note that launch promotions expire: every row on this site carries its price source and verification date, so budget long-term against the list rate.

GPT-5.5 or Claude Sonnet 5 — which is better?+

GPT-5.5 scores higher on multimodal and overall reasoning, ideal for flagship tasks mixing text, image and audio. Claude Sonnet 5 excels at coding and long-task reliability at a friendlier price. Favor Claude for code-heavy work and GPT-5.5 for an all-round flagship, then compare quality and cost on your actual tasks.

How is LLM API pricing calculated? What are input/output tokens?+

LLM APIs bill separately per million tokens for input (what you send) and output (what the model generates), with output usually several times more expensive. Monthly cost ≈ (avg input tokens × rate + avg output tokens × rate) × monthly requests. Use our cost calculator to estimate it directly.

Which LLM is best for RAG applications?+

RAG depends on both context window size and input token price, since retrieved documents consume a lot of input. Gemini's huge context plus low input pricing suits RAG well; for smaller corpora, a cheaper mid-tier model with good retrieval and caching keeps costs down.

How do I choose the right AI model?+

Define your quality bar first, then compare input/output pricing, latency and speed, context length and modalities. Start from a scenario (coding, RAG, production), narrow down with the comparison table and recommendation wizard, and validate your monthly budget with the cost calculator.

Explore more tools and content

From the model encyclopedia to rankings and guides, find everything you need.