LLM Insights中文
Back to all articles

LLM Pricing Changes, August 2026: DeepSeek Goes Time-of-Day, OpenAI Cuts Its Best Model to $4/$20, Ling-3.0 Enters at $0.07

11 min readLLM Price Compare

August 2026 did not look like the last two years of LLM pricing, where everybody cut rates in roughly the same direction at roughly the same time. Five things happened this month worth writing down, and only two of them were plain reductions in a number. The rest changed how billing works, which is the more consequential kind of change: a rate cut can be absorbed with multiplication, but a billing-model change invalidates the way you were estimating in the first place.

Every figure below comes from this site's catalog as recorded in August 2026, with verification dates noted. Promotional rates and list rates are kept separate throughout.

1. DeepSeek brought time-of-day pricing to LLM APIs

This is the month's biggest change and the least discussed. All three of DeepSeek's current models moved to two-tier billing: peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, billed at 2x list, with off-peak rates the rest of the time.

Essentially nobody else in this market does this. It is cloud-compute pricing logic — using price to shift load away from contended hours — imported wholesale into token billing.

ModelOff-peak inOff-peak outPeak inPeak outQuality
deepseek-v4-flash$0.14$0.28$0.28$0.5688
deepseek-chat-v4$0.2574$1.0287$0.5148$2.057488
deepseek-reasoner-v4$0.435$0.87$0.87$1.7492
DeepSeek two-tier pricing, per 1M tokens (verified 2026-08-16)

Where you are matters more than it used to. In UTC+8 those windows land at 09:00 to 12:00 and 14:00 to 18:00 local, covering almost the whole working day. A team in Taipei or Singapore running an internal tool on office-hours traffic is paying the doubled rate nearly all the time, not the one printed on the comparison table. In US Eastern time the peak windows fall overnight, so the same model costs meaningfully less for the same workload.

Read the other way, this is an unusual opportunity for schedulable work. Overnight batch summarisation, offline data cleaning, asynchronous generation — anything you can push off-peak buys deepseek-reasoner-v4 at $0.435/$0.87 with a quality score of 92. The other model scoring 92 in our catalog, claude-sonnet-5, charges $2/$10 — 11.5x more for output. What reasoner-v4 gives up is throughput: 30 tokens/second against Sonnet 5's 95, the slowest figure we track.

Two implementation notes. Express the schedule in UTC rather than deriving it from local time, because daylight-saving arithmetic on the client side is how an entire batch quietly runs at 2x. And keep a peak-hours fallback: qwen3.5 charges a flat $0.10/$0.15 around the clock, undercutting even DeepSeek's off-peak rate on both sides.

2. OpenAI cut its highest-scoring model to $4/$20

gpt-5.6-sol moved from $5/$30 to $4/$20 on 22 August against an official source — 20% off input, a third off output. Its quality score of 98 is the highest among the 68 models we track.

The size of the cut matters less than where it lands. Five models in the catalog score 96 or above: gpt-5.6-sol (98), gpt-5.5 (97, $5/$30), claude-opus-5 (97, $5/$25), and claude-fable-5 and claude-mythos-5 (both 96, $10/$50). After the cut, the highest-scoring model is also the cheapest one on both input and output. That has not been true of the flagship tier in the two years we have been tracking it.

claude-opus-5 absorbs the impact most directly. Before the cut it was 17% cheaper on output and the two had genuinely different cases. Now Sol is 20% cheaper on both sides, one point higher, matches the 1,049K context and adds native audio that Opus 5 does not have. Every column moved the same way at once.

Winning the table is not the same as being worth a migration. A composite quality score cannot see tool-use reliability, instruction adherence deep into a long conversation, or the fact that your prompts were tuned against one particular model. Switching vendors means a full regression pass, and in this tier that usually costs more than a year of the price difference. Start new projects on Sol; run existing systems against shadow traffic before committing.

3. Anthropic's flagship line: Opus 5 arrives, Sonnet 5 drops another third

claude-opus-5 entered our catalog in August at $5/$25, quality 97, 1,049K context. The number only means something against Anthropic's own trajectory: our price history records claude-opus-4-8 at $15/$75 in January 2026, $10/$50 in April and $5/$25 by June. Today's flagship rate is a third of what it was at the start of the year, while capability moved the other way.

In the same month claude-sonnet-5 went from July's $3/$15 to $2/$10, a third off both sides. The previous generation, claude-sonnet-4-6, is still listed at $3/$15 with quality 90 — so the newer tier is simultaneously 33% cheaper and two points better. That combination is rarer than it sounds. Vendors shipping a new tier usually hold one axis and improve the other.

claude-fable-5 is the case that needs explaining separately. It is $10/$50 at quality 96, while Opus 5 is $5/$25 at quality 97 — same vendor, same 1,049K context, effectively the same throughput, double the price, and not one column in its favour. The fable line is positioned around long-form writing and narrative consistency, which happens to be exactly what automated benchmark scoring measures worst. The honest position: the premium cannot be justified from a table, only from a blind test on your own material. Worth running if you produce prose at volume that reaches readers largely as generated. Otherwise baseline on Opus 5.

4. Gemini Flash: the list price did not move, the discount did

A correction first, because this one circulates in a distorted form. gemini-3.5-flash has listed at $1.50/$9.00 since it entered our catalog and did not change this month. Reports of a Gemini 3.5 Flash price cut are not describing the list rate.

What got cheaper is the set of billing options around it. Google offers Batch and Flex tiers at half list ($0.75/$4.50), and cache-hit input at $0.15 per 1M tokens, a tenth of the list rate. Which means what you pay depends far more on whether your workload tolerates asynchronous execution and repeated prefixes than on which model you selected.

No comparison table can see that distinction, and it frequently matters more than the choice of model. A pipeline that runs in Batch behind a stable system prompt can land between a tenth and a half of list. A synchronous one-shot request with a fresh prefix every time pays $1.50/$9.00.

Related: if you do not need the 1,049K context or the audio modality, Google's own gemini-2.5-flash is $0.30/$2.50 at quality 84. Against 3.5-flash's 86, that is five times the input rate and 3.6 times the output rate for two points. Worth doing the arithmetic before defaulting upward.

5. InclusionAI Ling-3.0-flash enters at $0.07

Ling-3.0-flash joined the catalog in August at $0.07/$0.22, quality 80, 262K context, 180 tokens/second. It is a 124B-parameter mixture-of-experts model activating roughly 5.1B parameters per token, which is the technical reason it reaches that speed at that price: inference cost tracks active parameters, not total ones.

One caveat has to be stated plainly. The list price is ¥0.40/¥1.20 per 1M tokens. The dollar figures above reflect a 65% launch promotion currently running on OpenRouter. Build an annual budget against list, not against the promotional rate, and make sure your architecture can swap the model underneath when the promotion ends.

Against the rest of the 180 tokens/second band: claude-haiku-4-5 is $1/$5 at quality 82 with 200K of context — roughly 14x the input rate and 23x the output rate for two points of quality and a smaller window. deepseek-v4-flash is $0.14/$0.28 at quality 88 with 1,000K context, better on both of those axes, but subject to the peak surcharge above. Ling has no time-of-day component at all.

Quality 80 draws the boundary: enough for classification, routing, extraction, rewriting and first-pass filtering; not enough to carry reasoning. The highest-value use is the top of a tiered stack, absorbing the bulk of traffic at near-zero marginal cost and escalating only hard cases. The question to ask of any ultra-cheap model is not whether it can replace your main one, but how much unnecessary traffic it can take away from it.

The August map: every model referenced above

ModelProviderInput /1MOutput /1MQualityContextSpeed
gpt-5.6-solOpenAI$4.00$20.00981,049K85
claude-opus-5Anthropic$5.00$25.00971,049K58
claude-fable-5Anthropic$10.00$50.00961,049K60
kimi-k3Moonshot AI$3.00$15.00951,000K60
claude-sonnet-5Anthropic$2.00$10.00921,049K95
deepseek-reasoner-v4DeepSeek$0.435$0.8792131K30
grok-4.5xAI$2.00$6.0091500K91
gemini-3.1-proGoogle$2.00$12.00911,049K70
qwen3.5Alibaba$0.10$0.15911,000K85
gpt-5.6-lunaOpenAI$0.20$1.20901,049K150
deepseek-chat-v4DeepSeek$0.2574$1.028788131K140
gemini-3.5-flashGoogle$1.50$9.00861,049K190
claude-haiku-4-5Anthropic$1.00$5.0082200K180
gpt-4o-miniOpenAI$0.15$0.6080128K160
Ling-3.0-flashInclusionAI$0.07$0.2280262K180
Speed in tokens/second. DeepSeek rates are off-peak; Ling-3.0-flash is promotional.

A few models in the table are not discussed above but have their own analysis pages:

  • deepseek-chat-v4 — the 131K context is its real constraint, with 262K to 1,000K available in the same price band.
  • gemini-3.1-pro — $2/$12 at quality 91, squeezed by claude-sonnet-5 and grok-4.5; the case for it is GCP, not rate.
  • kimi-k3 — a $3/$15 flagship, though gpt-5.6-terra matches its score of 95 while beating it on price, context and speed.
  • gpt-4o-mini — still running in production everywhere, with a price advantage much smaller than assumed.

Full price tables by provider: OpenAI, Anthropic, Google, DeepSeek, xAI, Alibaba, Mistral, Moonshot AI.

What the five changes add up to

The takeaway is not which model is cheapest. It is that a single price figure is losing its explanatory power. DeepSeek's rate depends on when you send the request. Google's depends on whether you can use Batch and caching. InclusionAI's depends on whether a promotion is still running. Anthropic and OpenAI both cut their flagships hard enough that the ordering between them flipped on 22 August. A single dollars-per-million-tokens column can no longer tell you what you will actually pay across those five vendors.

The other structural signal: a quality score of 90 or above no longer implies expensive. Twenty-five models in our catalog score 90+, and four of them price output below $1.20 — qwen3.5 at $0.15, MAI-DS-R1 at $0.28, deepseek-reasoner-v4 at $0.87 and gpt-5.6-luna at $1.20. The assumption that a cheap model comes with a visible capability ceiling did not survive August 2026.

The practical response is tiering. Let something in the Ling-3.0-flash or gpt-5.6-luna class absorb the bulk of traffic. Reserve claude-opus-5 or gpt-5.6-sol for work where failure is expensive and no human reviews the output. Give the middle to claude-sonnet-5 or grok-4.5.

Run the numbers on your own workload

Enter tokens per request and monthly volume to compare what these models actually cost you.

Open the cost calculator

Frequently asked questions

When exactly are DeepSeek's peak hours in my time zone?

Peak billing at 2x applies from 01:00 to 04:00 and 06:00 to 10:00 UTC. That is 09:00 to 12:00 and 14:00 to 18:00 in UTC+8, and 21:00 to 00:00 and 02:00 to 06:00 in US Eastern (UTC-5). Configure schedulers in UTC directly rather than converting from local time, since a daylight-saving error can push an entire batch into the doubled rate.

After the gpt-5.6-sol cut, is there still a reason to use Claude Opus 5?

Not on the numbers. gpt-5.6-sol moved to $4/$20 at quality 98 on 22 August, beating claude-opus-5's $5/$25 and quality 97 on input rate, output rate and score, with the same 1,049K context and audio support Opus 5 lacks. But a composite quality score does not measure tool-use reliability, instruction adherence across long conversations, or the sunk cost of prompts already tuned against one model. A cross-vendor migration needs a full regression pass, which in this tier often costs more than a year of the price difference. Start new projects on Sol and compare shadow traffic before moving existing ones.

Did Gemini 3.5 Flash actually get cheaper this month?

Its list price did not change and remains $1.50/$9.00 per 1M tokens. What changed is the billing options around it: Batch and Flex tiers run at half list ($0.75/$4.50), and cache-hit input is $0.15 per 1M tokens. Whether your effective cost fell depends entirely on whether your workload can use those modes.

Is Ling-3.0-flash's $0.07 rate permanent?

No. The list price is ¥0.40/¥1.20 per 1M tokens; $0.07/$0.22 reflects a 65% launch promotion on OpenRouter. Budget against the list rate for anything longer than a quarter, and confirm your abstraction layer can switch models once the promotion ends.

Why would anyone pick Claude Fable 5 over Opus 5 at twice the price and a lower score?

Our quality score is a composite benchmark figure that tracks automatically gradable properties such as task accuracy. The fable line targets long-form writing and narrative consistency, which is what automated grading measures least reliably. One point lower does not mean the prose is worse, and it does not guarantee it is better either — the difference can only be settled by a blind test on your own material. Worth testing if you publish generated prose at volume; otherwise baseline on Opus 5.

Which model scoring 90 or above is the cheapest?

By output rate, Alibaba's qwen3.5 at $0.10/$0.15 with quality 91 and a 1,000K context. Two caveats: our price for it comes from OpenRouter (verified 2026-08-06) rather than a vendor pricing page, and strategic pricing at that level can be revised. Confirm the official rate and keep the ability to switch models before committing to it long term.