LLM Insights中文

Compare LLM API prices and find the best-value AI model

Compare input/output pricing, context and quality across 30+ models — GPT, Claude, Gemini, DeepSeek and more — to pick the right API fast.

Featured Models

OpenAI logo
gpt-5.5
OpenAI
97

All-round flagship with top reasoning & multimodal skills

Input Price$5.00/1M tokens
Output Price$30.00/1M tokens
Context Window1049K
Output Speed85 tok/sMedium
Text
Vision
Audio
Anthropic logo
claude-opus-5
Anthropic
97

Anthropic's newest flagship — top coding & long-task quality

Input Price$5.00/1M tokens
Output Price$6.25/1M tokens
Context Window1049K
Output Speed58 tok/sMedium
Text
Vision
Google logo
gemini-3.5-flash
Google
86

Huge context, fast and cheap multimodal model

Input Price$1.50/1M tokens
Output Price$9.00/1M tokens
Context Window1049K
Output Speed190 tok/sFast
Text
Vision
Audio
xAI logo
grok-4.3
xAI
89

New-gen model blending real-time knowledge & reasoning

Input Price$2.00/1M tokens
Output Price$6.00/1M tokens
Context Window131K
Output Speed100 tok/sFast
Text
Vision
DeepSeek logo
deepseek-chat-v4
DeepSeek
88

Open-source value king with rock-bottom cost

Input Price$0.26/1M tokens
Output Price$1.03/1M tokens
Context Window131K
Output Speed140 tok/sFast
Text

Quick price comparison

Sorted by value by default — see who's most worth it at a glance.

ProviderModel NameTry
Microsoft logoMicrosoft
Phi-4-mini
ChatbotAffordableFast
68
250tok/sFast
$0.02$0.041700.0128KTry
01.AI logo01.AI
Yi-Lightning
ChatbotAffordableFast
78
200tok/sFast
$0.14$0.14557.116KTry
Microsoft logoMicrosoft
Phi-4
CodingChatbotAffordableFast
76
180tok/sFast
$0.07$0.14542.916KTry
Baichuan logoBaichuan
Baichuan4-Air
ChatbotAffordableFast
70
180tok/sFast
$0.14$0.14514.732KTry
Microsoft logoMicrosoft
MAI-DS-R1
ReasoningCoding
90
60tok/sMedium
$0.14$0.28321.4128KTry
Microsoft logoMicrosoft
Llama-3.3-70B-Instruct
ChatbotCoding
78
100tok/sFast
$0.10$0.32243.8128KTry
OpenAI logoOpenAI
gpt-4o-mini
ChatbotAffordableFast
80
160tok/sFast
$0.15$0.60133.3128KTry
Alibaba logoAlibaba
qwen-plus
ChatbotCodingAffordable
84
110tok/sFast
$0.26$0.78107.7131KTry
DeepSeek logoDeepSeek
deepseek-reasoner-v4
ReasoningCoding
92
30tok/sSlow
$0.43$0.87105.7131KTry
Groq logoGroq
llama-3.3-70b-versatile
ChatbotAffordableFast
78
300tok/sFast
$0.59$0.7998.7128KTry
Together AI logoTogether AI
meta-llama/Llama-3.3-70B-Instruct-Turbo
ChatbotAffordable
78
110tok/sFast
$0.88$0.8888.6131KTry
DeepSeek logoDeepSeek
deepseek-chat-v4
ChatbotAffordableFast
88
140tok/sFast
$0.26$1.0385.5131KTry
Baidu logoBaidu
ERNIE X1
ReasoningAffordable
86
45tok/sSlow
$0.28$1.1078.2128KTry
Together AI logoTogether AI
Qwen/Qwen2.5-72B-Instruct-Turbo
ChatbotCodingAffordable
79
100tok/sFast
$1.20$1.2065.8131KTry
Anthropic logoAnthropic
claude-haiku-4-5
ChatbotAffordableFast
82
180tok/sFast
$1.00$1.2565.6200KTry
Baidu logoBaidu
ERNIE 4.5
ChatbotMultimodal
84
90tok/sFast
$0.55$2.2038.2128KTry
Baidu logoBaidu
ERNIE 5.1
ChatbotReasoning
90
80tok/sFast
$0.59$2.6534.0128KTry
OpenAI logoOpenAI
o4-mini
ReasoningCodingAffordable
88
80tok/sMedium
$1.10$4.4020.0200KTry
xAI logoxAI
grok-4.3
ChatbotMultimodal
89
100tok/sFast
$2.00$6.0014.8131KTry
Mistral logoMistral
mistral-large-3
ChatbotCoding
82
80tok/sMedium
$2.00$6.0013.7128KTry
Microsoft logoMicrosoft
Mistral-Large
ChatbotCoding
82
80tok/sMedium
$2.00$6.0013.7128KTry
OpenAI logoOpenAI
o3
ReasoningCoding
95
40tok/sSlow
$2.00$8.0011.9200KTry
Mistral logoMistral
mistral-medium-3.5
ChatbotCoding
86
90tok/sFast
$1.50$7.5011.5128KTry
OpenAI logoOpenAI
gpt-4o
ChatbotMultimodal
88
100tok/sFast
$2.50$10.008.8128KTry
Cohere logoCohere
command-r
ChatbotAffordable
75
120tok/sFast
$2.50$10.007.5128KTry
Baichuan logoBaichuan
Baichuan4
ChatbotReasoning
80
60tok/sMedium
$13.89$13.895.832KTry
Guided picks

Choose a model by your scenario

Not sure where to start? Jump in from the three most common needs.

What's new

Recently added and updated models

Prices and models are refreshed continuously by our daily data pipeline.

View all models
FAQ

Questions you might have about LLM API pricing

What's the cheapest LLM API in 2026?

By per-million-token pricing, open-source models like DeepSeek, plus Gemini Flash Lite and GLM, are usually the cheapest — output prices can be 10× lower than flagship models. The best choice depends on your quality bar, so sort the table above by price and cross-check the value score.

GPT-5.5 or Claude Sonnet 5 — which is better?

GPT-5.5 scores higher on multimodal and overall reasoning, ideal for flagship tasks mixing text, image and audio. Claude Sonnet 5 excels at coding and long-task reliability at a friendlier price. Favor Claude for code-heavy work and GPT-5.5 for an all-round flagship, then compare quality and cost on your actual tasks.

How is LLM API pricing calculated? What are input/output tokens?

LLM APIs bill separately per million tokens for input (what you send) and output (what the model generates), with output usually several times more expensive. Monthly cost ≈ (avg input tokens × rate + avg output tokens × rate) × monthly requests. Use our cost calculator to estimate it directly.

Which LLM is best for RAG applications?

RAG depends on both context window size and input token price, since retrieved documents consume a lot of input. Gemini's huge context plus low input pricing suits RAG well; for smaller corpora, a cheaper mid-tier model with good retrieval and caching keeps costs down.

How do I choose the right AI model?

Define your quality bar first, then compare input/output pricing, latency and speed, context length and modalities. Start from a scenario (coding, RAG, production), narrow down with the comparison table and recommendation wizard, and validate your monthly budget with the cost calculator.

Explore more tools and content

From the model encyclopedia to rankings and guides, find everything you need.