If you only remember three numbers, make them "70x, 180x, 2%." That's roughly the input-price gap and output-price gap between the cheapest and most expensive models in this comparison, and how low a caching-hit rate can drop relative to the uncached price. Together, those numbers sketch the real picture of AI model pricing in August 2026: it's no longer a simple "more expensive = smarter" ladder — it's a genuinely complicated ledger built from caching discounts, long-context tiering, peak-hour surge pricing, reasoning-effort switches, and subscription-plus-metered billing running side by side. And that ledger moves faster than most people assume: within our research window alone, Claude Opus 5 quietly took over from Opus 4.8, Qwen3.8-Max went from an unpriced preview to a fully priced GA release overnight, DeepSeek was reported to be rolling out the only time-of-day surge pricing scheme in this entire comparison, and Google's much-anticipated Gemini 3.5 Pro slipped from June to July to, now, "no date given." This piece lays that ledger out in full.
This piece covers current flagship models from OpenAI, Anthropic, Google, DeepSeek, Moonshot (Kimi), Alibaba (Qwen), Zhipu (GLM), xAI, and MiniMax. Our method has two layers: first, we reuse pricing figures already cross-verified in this site's own reviews (DeepSeek V4 Flash, Grok 4.5 cost-performance, GLM-5.2, Kimi K3 vs. Qwen3.8-Max-Preview, the OpenAI Astra preview, MiniMax H3, and others); second, we did fresh, independent research specifically on the fastest-moving items this round, and we flag anything sources disagree on as "unverified" rather than inventing a clean number. Before publishing, we also ran an extra round of verification specifically on DeepSeek's rumored peak-hour surge pricing and Qwen3.8-Max's newly-GA pricing — what that check turned up is called out explicitly below.
Contents
- 1. The full pricing table: nine labs' current lineups at a glance
- 2. OpenAI — GPT-5.6's three tiers just got cheaper, Astra is still a blank
- 3. Anthropic — Opus 5 swaps in at the same price, Fable/Mythos 5 are the ceiling
- 4. Google — the flagship actually on sale is 3.1 Pro; 3.5 Pro still isn't out
- 5. DeepSeek — from a permanent 75% cut to peak-hour surge pricing (verified)
- 6. Moonshot/Kimi — K3's rate doesn't move, but thinking is locked to max
- 7. Alibaba/Qwen — Qwen3.8-Max just went GA (verified)
- 8. Zhipu/GLM — the official price isn't even the cheapest place to buy it
- 9. xAI — Grok 4.5 is live, 4.6/4.7 exist only in tweets
- 10. MiniMax — M3 is the text/agent flagship, don't confuse it with H3
- 11. What actually decides your bill: caching, surge pricing, reasoning effort, speed tiers
- 12. The relay-provider view: open-weight and closed models are two different economies
- 13. Six takeaways
- 14. Data freshness and what we re-verified before publishing
- 15. Conclusion: how to actually use this table
1. The full pricing table: nine labs' current lineups at a glance
Let's start with the numbers themselves. The table below groups models by vendor, roughly sorted by input price, covering everything currently GA (generally available) across all nine labs as of early August 2026. Where sources genuinely disagree, we write "unverified" rather than forcing a single number.
| Vendor | Model | Status | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|---|---|
| DeepSeek | V4-Flash | GA (0731 release) | $0.14 (cache hit $0.0028) | $0.28 | 1M | No peak surcharge; among the cheapest input prices anywhere |
| Alibaba/Qwen | Qwen3.6-35B-A3B | GA | $0.14 | $0.90 | — | Cheapest multimodal model in this comparison |
| DeepSeek | V4-Pro | GA | $0.435 (off-peak) | $0.87 (off-peak) | 1M | A 2x peak-hour surcharge has been announced but, as of publication, is not yet active — see section 5 |
| MiniMax | M3 (text/agent model) | GA (June 1) | $0.30 (list $0.60, "permanent 50% off") | $1.20 | 1M (doubles past 512K) | 428B total / 23B active MoE, native multimodal + computer use |
| MiniMax | M2.1 (prior gen) | GA | $0.30 | $1.20 | ~262K | Same price as M3, which mainly sells context length and speed |
| Zhipu/GLM | GLM-5.2 | GA (June 13, MIT-licensed weights) | $1.40 (cache hit $0.26) | $4.40 | 1M (128K max output) | Official price vs. ~$0.55/$1.85 third-party median — see sections 8 and 12 |
| Qwen | Qwen3.8-Max | Just went GA August 3 | $2.00 | $6.00 | 1M (65,536 max output) | 2.4T total / ~95B active MoE; open weights due "next week"; re-verified below |
| xAI | Grok 4.5 | GA (July 8) | $2.00 (cache hit $0.50) | $6.00 | 500K (down from prior gen's 1M) | Best cost-performance among closed multimodal flagships, not the cheapest overall |
| OpenAI | GPT-5.6 Luna | GA (80% price cut July 30) | $0.20 (cache write $0.25 / read $0.02) | $1.20 | 1.05M (tiers past 272K) | Long-context tier: $0.40/$1.80 |
| Moonshot | Kimi K2.7 (prior gen) | GA | $0.95 (cache hit $0.19) | $4.00 | 262,144 | Cheaper than K3 but a much smaller context window |
| Anthropic | Claude Sonnet 5 | GA | $3.00 (promo $2.00 through 2026-08-31) | $15.00 (promo $10.00) | 1M | Reverts to standard price after the promo; supports 2576px vision |
| OpenAI | GPT-5.6 Terra | GA (20% price cut July 30) | $2.00 (cache write $2.50 / read $0.20) | $12.00 | 1.05M | Long-context tier: $4.00/$18.00 |
| Anthropic | Claude Opus 5 (replaces Opus 4.8) | GA (July 24) | $5.00 | $25.00 | 1M (128K max output) | Same price as Opus 4.8; adds a low/mid/high "effort" switch |
| OpenAI | GPT-5.6 Sol | GA | $5.00 (cache write $6.25) | $30.00 | 1.05M | "Fast mode": 2x price for 2.5x speed |
| Moonshot | Kimi K3 | GA (July 16) | $3.00 (cache hit $0.30) | $15.00 | 1M (flat, no long-context tier) | Thinking mode locked to max, can't be disabled; subscription tiers also exist |
| Anthropic | Claude Fable 5 | GA | $10.00 | $50.00 | 1M (default is the max) | Thinking is always on; requires 30-day data retention |
| Anthropic | Claude Mythos 5 | GA but restricted access | $10.00 | $50.00 (down sharply from a restricted-preview $25/$125) | Shares Fable's base | Only for vetted "Project Glasswing" partners and specific biosecurity researchers, no self-serve access |
| Gemini 3.5 Flash-Lite | GA | $0.30 | $2.50 | — | Lightweight tier | |
| Gemini 3.6 Flash | GA | $1.50 | $7.50 | — | Claimed up to 65% lower cost on long-horizon agent tasks | |
| Gemini 3.5 Flash | GA | Third-party figures conflict ($0.15–$1.50) | ~$9.00 (also unverified) | — | We don't adopt a single number given the disagreement | |
| Gemini 3.1 Pro (top tier actually on sale) | GA (Preview) | $2.00 (≤200K) / $4.00 (>200K) | $12.00 (≤200K) / $18.00 (>200K) | 200K tier | Some sources claim a recent drop to $1/$6, unconfirmed against official docs |
A few names exist only as teasers or rumors right now, with zero official pricing. We list them for completeness, but every column below is genuinely "undisclosed":
| Vendor | Model | Status | Pricing | Notes |
|---|---|---|---|---|
| Gemini 3.5 Pro | Not released ("in partner testing"; GA date slipped from June to July to "TBD") | Undisclosed | Rumored 2M-token context, unconfirmed | |
| Gemini 4 | Training, not released | Undisclosed | Pichai said July 26 it's "much bigger than before"; historical cadence suggests Nov–Dec | |
| xAI | Grok 4.6 | Not released (Musk says ~August 7) | Undisclosed | ~1.5T params, Musk's own words only, no official confirmation |
| xAI | Grok 4.7 | Not released ("weeks" after 4.6) | Undisclosed | ~2.1T params, same caveat |
| OpenAI | Astra | Teased only, not released | Undisclosed | The "$2,000" cost estimate uses existing Sol pricing as a reference point, not Astra's own rate |
A quick note on terminology: per-token means straightforward metered billing; time-of-day pricing means the rate floats within the same pricing scheme depending on the hour; subscription-plus-metered means a monthly plan and a pay-as-you-go API coexist as two separate billing systems that don't convert cleanly into each other — a point our earlier Kimi K3 vs. Qwen3.8-Max-Preview review covers in detail. Now, vendor by vendor.
2. OpenAI — GPT-5.6's three tiers just got cheaper, Astra is still a blank
Sol, Terra, and Luna are OpenAI's only current lineup, and they just got a substantial July 30 price cut (Terra down 20%, Luna down 80%, Sol unchanged) — widely read as a direct response to enterprise doubts about AI spend ROI and pricing pressure from cheap Chinese open-weight models. All three tiers share a 1.05M context window and 128K max output, but all three step into a long-context tier once input passes 272,000 tokens (roughly doubling): Luna's $0.20/$1.20 becomes $0.40/$1.80, Terra's $2/$12 becomes $4/$18, and Sol's $5/$30 becomes $10/$45.
Sol also has a unique "Fast mode": 2x the price for 2.5x the speed — the only case in this whole comparison where a vendor puts an explicit price tag on response speed itself; everyone else's "faster" is either free or simply not offered as a choice. As for the next-gen model Astra, we've already covered it in depth on this site: it's the codename dropped almost in passing in an August 1 research blog post about math results, and as of publication OpenAI still hasn't disclosed a release date, model size, context window, or API pricing — the widely-cited "$2,000 in compute" figure also uses existing Sol pricing as a stand-in, not Astra's own rate. In plain terms: if you're trying to budget for Astra's pricing today, the honest answer is that it's too early.
3. Anthropic — Opus 5 swaps in at the same price, Fable/Mythos 5 are the ceiling
One change from this cycle that's easy to miss but worth calling out on its own: Claude Opus 5 launched July 24 and replaced Opus 4.8 as the current flagship, positioned officially as approaching Fable 5's frontier intelligence "at half the price," while keeping the exact same pricing as Opus 4.8 ($5/$25). If you still have an old plan built around "Claude Opus 4.8" as a model name, it should now read as "replaced by a same-priced successor," not "still the latest flagship." Opus 5 adds a low/mid/high "effort" switch, but this is not a separate price tier — it doesn't change the unit rate, it changes how much reasoning (and therefore how many tokens) gets consumed per request, so the rate card itself still has just one number.
Sonnet 5 currently carries a promotional rate of $2/$10 (through 2026-08-31), reverting to a standard $3/$15 after — if you're reading this past that date, verify whether it's reverted. Fable 5 ($10/$50) is positioned away from cost-performance entirely: thinking is always on and can't be turned off, and enterprise customers must accept 30-day data retention (no zero-data-retention option). Mythos 5 (also $10/$50, sharing Fable's base but with lighter safety guardrails) has come down sharply from its restricted-preview price of $25/$125 back in April, but it's still not available through the public API, claude.ai, or any self-serve channel — it's limited to vetted "Project Glasswing" partners and specific biosecurity researchers, so regular developers can't access it regardless of price.
4. Google — the flagship actually on sale is 3.1 Pro; 3.5 Pro still isn't out
Worth clearing up a common point of confusion first: the Gemini 3.5 Pro Google teased at May's I/O still hasn't launched as of August 5 — the GA date has slipped from June to July, and as of July 21 the official line was still "in partner testing" with no new date given. There's currently no official pricing and no official context window figure (a rumored 2M tokens remains unconfirmed). If you see anyone online claiming "Gemini 3.5 Pro is live and here's the price," treat it as false.
The actual top tier currently for sale is Gemini 3.1 Pro (Preview): $2/$12 up to 200K context, $4/$18 beyond it. A few sources claim a recent drop to $1/$6, but that's not confirmed against official docs, so we don't treat it as settled. On the lighter end, Gemini 3.6 Flash ($1.50/$7.50) and Gemini 3.5 Flash-Lite ($0.30/$2.50) are both confirmed; but Gemini 3.5 Flash's pricing still can't be cross-verified this round — sources conflict wildly, with input quoted anywhere from $0.15 to $1.50 — the exact same issue this site's DeepSeek V4 Flash review already flagged, and we again decline to adopt a single number just to fill a cell. Gemini 4 remains in training; Pichai said in a July 26 interview it's "much bigger than before," and historical release cadence points to a November–December launch, with no pricing signal at all. One more note: Google's Batch API is 50% off across the board with a 24-hour turnaround, and cached input runs at roughly 10% of the uncached rate.
5. DeepSeek — from a permanent 75% cut to peak-hour surge pricing (verified)
DeepSeek V4-Flash ($0.14/$0.28) moved from preview to the official "0731" GA build on July 31 — same architecture, but retrained enough to substantially lift agentic benchmark scores. Pricing didn't change, and there's no peak surcharge, which is itself a selling point over V4-Pro's more predictable-in-theory-but-actually-less-predictable billing. The bigger story is V4-Pro: this round's research surfaced a change our earlier site coverage hadn't caught — multiple outlets reported that, weeks after DeepSeek's "permanent 75% price cut," the company turned around and announced a 2x peak-hour surcharge for 9:00–12:00 and 14:00–18:00 Beijing time on V4-Pro. That would make DeepSeek the only major vendor in this entire comparison using time-of-day pricing — a sharp contrast with Kimi K3's flat rate regardless of the clock.
Verified before publishing: DeepSeek's peak-hour pricing still isn't live, and it would apply to input and output alike
Since the original research flagged this as possibly stale, we ran an independent check specifically on this claim before publishing. Two things came back confirmed. First: as of early August 2026, DeepSeek's official pricing page still shows a single flat rate, and multiple third-party pricing trackers that cross-reference the official docs explicitly state that "the peak-hour surcharge has been announced, but no effective date has been given, and it is not yet active" — so as of this writing, you're still being billed the off-peak rate, and there's no need to worry about your bill doubling today. Second, and this refines what the original research could confirm: this 2x surcharge is described by multiple sources as applying to all billing items — cache-hit input, cache-miss input, and output alike (e.g., off-peak output of ¥6/M would become ¥12/M at peak, with input scaling the same way) — rather than "output only," which is what earlier sourcing had been able to confirm. Once this mechanism actually goes live, check DeepSeek's official pricing page (api-docs.deepseek.com) directly.
On caching, DeepSeek is in a league of its own: V4-Flash's cache-hit rate is $0.0028 (about 98% off), and V4-Pro's is roughly $0.0036 — the steepest caching discount of any vendor in this comparison, and a real win for anyone repeatedly hitting the same long document or codebase.
6. Moonshot/Kimi — K3's rate doesn't move, but thinking is locked to max
Kimi K3 ($3/$15, cache hit $0.30, 90% off) launched July 16 and opened its full weights on July 27 (a 2.8T-parameter MoE, weight files around 1.4TB after MXFP4 quantization — one of the largest open-weight releases in history at the time), under a modified MIT license for commercial use. It charges a single flat rate across its full context window (up to 1M), unlike OpenAI or Google's "more expensive past X tokens" tiering — bill predictability is part of its pitch.
One thing worth flagging on its own: K3's thinking mode is locked to its highest intensity by default and can't be turned off. The unit rate doesn't change, but real output-token consumption is inherently higher, so your actual bill can run above what the headline rate implies — a markup delivered through hidden consumption rather than a raised sticker price, worth watching closely if you're routing through a relay provider. Due to a demand surge, Moonshot paused new-user signups on July 19 (reopening in batches); whether that's fully lifted by the time you read this is worth checking directly. The prior-gen Kimi K2.7 (the Code variant, $0.95/$4, cache hit $0.19, 262,144 context) is still available and a reasonable pick if you're more budget-constrained and don't need million-token context. Separately, a subscription tier (Moderato $19/mo, Allegretto $39/mo, Allegro $99/mo, Vivace $199/mo) runs alongside the metered API as a fully separate billing system — the Allegro/Vivace tiers unlock K3's million-token extended conversations.
7. Alibaba/Qwen — Qwen3.8-Max just went GA (verified)
A significant status change happened during this research window: Qwen3.8-Max officially launched on August 3 (it was previously "Preview" only, with no standalone pricing page — accessible only through the Token Plan subscription or the Qoder/QoderWork channels, with third parties left to reverse-engineer an estimated price). Post-launch, it now carries official pricing of $2.00/M input, $6.00/M output, a flat rate across its full 1M context window with no long-text surcharge, and it's a 2.4-trillion-parameter (roughly 95B active) MoE — Alibaba's largest and most capable flagship to date. Open weights are planned for release "next week" (within a week of launch).
Verified before publishing: Qwen3.8-Max's pricing hasn't changed, and it cross-checks from a second angle
Given this was a status change that happened mid-research-window, we ran a separate check before publishing. Result: the $2.00/$6.00 figures are still accurate as of publication, unchanged. We also found an independent way to cross-check that number — multiple sources describe Qwen3.8-Max's international pricing benchmark as "roughly 40% of Claude Opus 5's input price and 24% of its output price." Opus 5 is $5/$25; 40% and 24% of those work out to exactly $2.00 and $6.00 — two independent paths landing on the same answer, which gives us more confidence in the figure. The "open weights next week" plan also picked up additional corroborating sources in this check, with the direction unchanged.
This means the conclusion in our earlier "Kimi K3 vs. Qwen3.8-Max-Preview" review — that the comparison "genuinely couldn't be run on price" — no longer holds as of August 3: both now carry official $/M pricing, so a direct comparison is possible. Kimi K3's input price ($3) is 1.5x Qwen3.8-Max's ($2); its output price ($15) is 2.5x Qwen3.8-Max's ($6). Separately, Qwen3.6-35B-A3B ($0.14/$0.90), a lighter multimodal model, remains available and holds the title of cheapest multimodal model in this comparison.
8. Zhipu/GLM — the official price isn't even the cheapest place to buy it
GLM-5.2 ($1.40/$4.40, cache hit $0.26, with cache storage currently free for a limited time) debuted as a preview on June 13 and had its weights fully open-sourced under MIT three days later — 1M context window, 128K max output. It's the clearest case in this entire comparison of "the official price isn't the lowest price": third-party hosts (Together AI, Fireworks, DeepInfra, SiliconFlow, OpenRouter, and others) offer a median price around $0.55/$1.85, with rates on OpenRouter as low as $0.29/M input — roughly 20–50% of the official rate. We break down the mechanism behind this in section 12. Separately, the GLM Coding Plan subscription (positioned against Claude Code) runs Lite at an official $18/month (promotional pricing as low as ~$12.60/month), Pro at $72/month, and Max at $160/month, with peak-hour quota consumed at 3x the rate and off-peak at 2x.
9. xAI — Grok 4.5 is live, 4.6/4.7 exist only in tweets
Grok 4.5 ($2/$6, cache hit $0.50, 75% off) launched July 8 with an official context window of 500K tokens — notably smaller than the prior generation's 1M, the only case in this comparison where the context window actually shrank between generations. It steps into a higher pricing tier past 200,000 tokens, though the exact markup hasn't been disclosed consistently. It's the best cost-performance option among closed multimodal flagships, though not the cheapest overall. Grok 4.6 (~1.5T parameters) and Grok 4.7 (~2.1T parameters) are both unreleased — they exist only as public statements from Musk on X between July 24–28 (4.6 around August 7, 4.7 "a few weeks" after that), not as official xAI confirmation, and pricing and context windows are completely blank. One aside: xAI's voice model Grok Voice Think Fast 2.0 updated August 1, priced at $0.08 per minute of audio — a completely separate billing system from text-token pricing.
10. MiniMax — M3 is the text/agent flagship, don't confuse it with H3
Worth scoping this correctly first: MiniMax also released a video-generation model, H3 ($0.13–0.14/second, 2K resolution), around the same time — but that's a video model, outside this piece's scope of "high-performance text/reasoning models," and we mention it only to avoid confusion. The text/agent flagship is M3 (launched June 1): a 428B-total/~23B-active MoE using MSA (MiniMax Sparse Attention), a 1M-token context window, and native support for image/video input plus desktop computer-use. It's priced at $0.30/M input, $1.20/M output — the list price was originally $0.60/$2.40, and the "permanent 50% off" promotion effectively made $0.30/$1.20 the standing rate, matching the prior-gen M2.1; any single request past 512K tokens is billed at 2x the rate for that turn, and cached reads run $0.06/M. Among China's open-weight lineup, MiniMax's M-series sits in the low-to-mid price range but stands out for context length and multimodal capability, and it's already reachable through third-party inference platforms like SiliconFlow and WaveSpeed.
11. What actually decides your bill: caching, surge pricing, reasoning effort, speed tiers
Beyond the headline per-token rates sits a layer that's easy to overlook but potentially matters more for your actual bill: how each vendor handles caching discounts, long-context surcharges, time-of-day pricing, reasoning effort, and response speed varies enormously. Here's the roundup:
| Mechanism | Who uses it | Details |
|---|---|---|
| Cache-hit discount | Nearly universal, but the size of the discount varies a lot | DeepSeek is the most aggressive: cache-hit rates land around 2% of the uncached price; OpenAI's tiers run about 10% of uncached; Google's cached input is about 10% of uncached; Kimi K3 is about 10% (90% off); GLM-5.2 is about 19% (roughly 81% off); Grok 4.5 is 25% (75% off); Anthropic hasn't disclosed a crossverifiable official cache-discount ratio, so we don't fabricate one |
| Long-context tiering | OpenAI (272,000-token threshold across all three tiers), Google's Gemini 3.1 Pro (200K threshold), Grok 4.5 (200K threshold, markup undisclosed) | All "the longer the input, the higher the rate" schemes, in contrast to DeepSeek and Kimi K3's flat pricing regardless of length |
| Time-of-day (peak/off-peak) pricing | DeepSeek V4-Pro only (not yet live as of publication) | A planned 2x multiplier on all billing items (input and output alike) during 9–12 and 14–18 Beijing time, trading bill predictability for demand smoothing |
| Speed-tier pricing | OpenAI GPT-5.6 Sol's "Fast mode" (unique) | 2x price for 2.5x speed — the only vendor putting an explicit price on response speed itself |
| Reasoning-effort switches (don't change the rate, only consumption) | Anthropic Opus 5's low/mid/high "effort"; DeepSeek V4-Pro's Non-Think/Think High/Think Max; GLM-5.2's High/Max | Same rate, but heavier reasoning burns more output tokens — the bill goes up through consumption, not price |
| Thinking forced always-on, can't be disabled | Kimi K3 (locked to max), Claude Fable 5 (thinking is always on) | No option to "turn off thinking to save money" — the rate is unchanged, but actual spend is inherently higher |
Working through this table leads to one conclusion: the only two vendors genuinely treating "reasoning" itself as an independent, separately priced dimension are OpenAI (speed tiering) and DeepSeek (time-of-day tiering); everyone else's "reasoning intensity" toggle leaves the rate untouched and works through token consumption instead — the variable most likely to be overlooked in a head-to-head comparison, and potentially the one with the biggest real-world impact on spend. In other words, judging purely by "price per million tokens" is likely to understate your real bill.
12. The relay-provider view: open-weight and closed models are two different economies
This is the angle our core readers care about most. We split it into two cases, because the underlying price-formation mechanics are genuinely different.
12.1 Open-weight models: third-party hosts compete on price, and the official rate isn't the floor
DeepSeek's V4 line, GLM-5.2, Kimi K3, Qwen3.6/3.8-Max, and MiniMax M3 are all open-weight, meaning any third party with enough compute can self-host and resell inference — which creates a dynamic that simply doesn't exist for closed models: multiple hosting providers competing on price, so the official rate stops being the market floor. The clearest example this round is GLM-5.2: official pricing (Z.ai/bigmodel.cn) is $1.40/$4.40, but third-party hosts (Together AI, Fireworks, DeepInfra, SiliconFlow, OpenRouter, and others) sit at a median of roughly $0.55/$1.85, with rates on OpenRouter as low as $0.29/M input — about 20–50% of the official price. DeepSeek V4 and Kimi K3 both show similar multi-host price competition (Together, Fireworks, Modal, SiliconFlow, OpenRouter, and more), though we didn't find a citable, specific price gap for those as clean as GLM-5.2's. This "relay/aggregator prices below official" phenomenon is, at bottom, compute-market competition made possible by open weights — a mechanism that simply doesn't exist for closed models.
12.2 Closed models: relay-price gaps come from FX arbitrage, bulk purchasing, and reverse-engineered access — not compute competition
GPT, Claude, Gemini, and Grok are all closed, so no third party can self-host them — a relay provider can only resell official account quota at a markup, and that gap comes from a completely different place than it does for open-weight models: bulk-purchased discount quota resold at markup, custom "¥-to-$" conversion rates (effectively a disguised discount), subscription slicing, enterprise cloud forwarding, and grey-area channels like reverse-engineered web protocols or free IDE quota. Our own earlier provider reviews have documented plenty of concrete examples: TokenRiver settles at "1¥ = $1" (roughly 14% of official pricing — a figure from earlier published content, not re-verified this round, so check the provider's current page); AIGCBest runs a roughly 1.5¥/USD rate; AZAPI's Claude rate is roughly ¥2.5/USD (about 34% of official); and BLTCY runs 1–2.5¥/USD.
Fresh, independent research this round confirms the same pattern is still very much alive: as of an April 2026 third-party review, one platform's internal settlement rate was ¥2.4/$ (about a 67% discount off the standard conversion), another was ¥2/$ (about 72% off), and a third offered GPT-5.1-codex at a limited-time rate roughly 15% below official. The price gap on closed models isn't cheaper compute — it's cheaper account quota, and the size, stability, and account risk (silent downgrades, providers disappearing) of that gap are meaningfully different from open-weight hosting competition, which is why the two types of relay providers shouldn't be evaluated by the same logic. The one-line summary: comparing "official price" is only a starting point for open-weight models, but nearly the end point for closed ones — because a closed model's official price is the ceiling; a relay can only resell it at a discount, never reprice the underlying compute itself.
13. Six takeaways
Put the pricing, the mechanics, and the relay-provider ecosystem together, and a handful of patterns matter more than any individual number:
- The open/open-weight camp dominates raw per-token pricing across the board: DeepSeek V4-Flash ($0.14/$0.28) and Qwen3.6-35B-A3B ($0.14/$0.90) sit permanently at the bottom of the input-price range, and against the most expensive closed models (Claude Fable 5/Mythos 5 at $10/$50), the input gap runs roughly 70x and the output gap roughly 180x — that's not a discount, that's an order of magnitude.
- "Most expensive" doesn't mean "furthest behind" — closed labs are fighting a price war too: OpenAI's steep July 30 cuts to Terra/Luna, and Anthropic positioning Opus 5 as "Fable 5-level intelligence at half the price," are both direct responses to open-model pricing pressure — but closed-lab discounting mostly hits mid-tier models. Flagship tiers (Sol, Fable 5, Opus 5) haven't budged, meaning vendors still want a premium reserved for "the strongest model," and the price war is concentrated in the middle of the market.
- What you actually pay for an open-weight model depends on where you buy it, not just the sticker price: GLM-5.2's official price and its third-party hosting median differ by 2–5x, a phenomenon that simply doesn't exist for closed models, where the official price is the ceiling.
- Reasoning intensity mostly hits your bill through consumption, not the unit rate: outside OpenAI's speed tiering and DeepSeek's peak pricing, every other vendor's reasoning-effort toggle leaves the per-token rate untouched; Kimi K3 and Claude Fable 5 even force thinking always-on. For relay users, that means "price per million tokens" alone can understate real spend — what actually decides the bill is how many tokens a model burns to finish the same task.
- Caching discounts are an underrated lever, and vendors vary enormously here: DeepSeek's cache-hit rate is nearly zero (about 98% off), OpenAI/Google/Kimi K3 cluster around 90% off, GLM-5.2 is roughly 81% off, and Grok 4.5 is 75% off — for repeatedly processing the same long document or codebase in an agentic workflow, the caching discount can matter more to your real bill than the uncached input price itself.
- The pricing vacuum between "teased" and "officially launched" is where misinformation spreads easiest: Astra (OpenAI), Gemini 3.5 Pro/Gemini 4 (Google), and Grok 4.6/4.7 (xAI) all sit in a stage where the name or specs are known but pricing, context, and release date are a complete blank — any claim that "a relay provider already supports it" or cites a specific price should be treated as false, since we found no official source backing that up. The counterexample is Qwen3.8-Max, which flipped from "unpriced preview" to "GA with official $/M pricing" during this exact research window — a reminder that this kind of pricing vacuum usually doesn't last long once a model actually goes GA.
One more pattern worth noting: subscription/credit-based billing coexisting with pure metered pricing is more common among Chinese vendors — the GLM Coding Plan, Kimi's membership tiers, and Qwen's Token Plan all run a monthly subscription alongside a metered API as two genuinely separate systems, and subscription pricing doesn't convert cleanly into an equivalent $/M-token rate. The US four (OpenAI/Anthropic/Google/xAI) still stick to pure metered pricing at the API layer; their consumer subscriptions (ChatGPT Plus/Pro, Claude Pro/Max, etc.) are a separate consumer-product layer that doesn't map directly onto developer API pricing, and the two shouldn't be conflated.
14. Data freshness and what we re-verified before publishing
Treat this table as a snapshot around August 5, 2026 — not a permanent rate card
Every price here is the result of independent, fresh research cross-checked against figures already verified on this site; anywhere marked "undisclosed/unverified" (all of Astra, all of Gemini 3.5 Pro/Gemini 4, all of Grok 4.6/4.7, Gemini 3.5 Flash's specific rate, Anthropic's specific caching-discount ratios, and whether DeepSeek V4-Pro's peak surcharge is officially live) reflects a genuine information gap or source conflict at research time — not "there will never be a number," just "no credible source had one as of publication." Promotional pricing (Claude Sonnet 5's $2/$10 promo through 2026-08-31, GLM Coding Plan's Lite-tier promo) has a hard expiration date; if you're reading this after that date, verify whether the standard rate has resumed. The relay-specific discount figures cited in section 12 (TokenRiver's "1¥=$1" and similar) carry over from this site's earlier published reviews and weren't individually re-verified this round — relay pricing typically moves faster than underlying model pricing, so check our own AI API relay comparison dashboard or the provider's current page directly.
Because pricing moves fast, we specifically re-verified the two most recent and most volatile claims in this round before publishing: DeepSeek V4-Pro's peak-hour pricing — confirmed still not live as of publication, with the official pricing page still showing a flat rate, and the planned surcharge covering all billing items (input and output alike), not just output; and Qwen3.8-Max's GA pricing — confirmed the $2.00/$6.00 figures are accurate, with an independent cross-check ("roughly 40%/24% of Claude Opus 5's rates") landing on the same numbers, which gives us added confidence. Both results are called out in the alert boxes in sections 5 and 7.
15. Conclusion: how to actually use this table
Having gone through all nine labs' pricing, the honest conclusion is that comparing "price per million tokens" alone is no longer enough to answer "which model is actually the better deal." A real decision needs at least three layers: the raw price itself, where the open-weight camp has an order-of-magnitude advantage; the billing mechanics — caching discounts, long-context tiering, peak pricing, reasoning effort — which can make the real bill several times higher or lower than the headline rate suggests; and, the layer our readers should care about most, where you're actually buying from — an open-weight model's official price is only a starting point, since third-party hosting competition can push it down to a fraction of that; a closed model's official price is the ceiling, and what a relay can offer is cheaper account quota, not cheaper compute, which is why the risk calculus for the two kinds of relay providers can't be applied interchangeably.
Put together, this means
If you just need something cheap and good enough, DeepSeek V4-Flash and Qwen3.6-35B-A3B are currently the price floor. If you need top-tier intelligence and budget is a secondary concern, Claude Opus 5, Fable 5, and GPT-5.6 Sol remain the ceiling. If you're optimizing for long-running, high-volume agentic workloads, caching-discount depth and whether "thinking" can be turned off may matter more than the headline rate.
- Budget-conscious developers optimizing for cost-performance → start with the open-weight camp (DeepSeek, Qwen, GLM, Kimi, MiniMax), and actually compare official vs. third-party hosting prices — don't assume the official rate is the floor.
- Teams needing top-tier intelligence with more budget headroom → Claude Opus 5, Fable 5, and GPT-5.6 Sol remain the current performance ceiling, but watch how effort/reasoning-intensity switches affect real consumption.
- Developers relying on a relay to reach overseas models → first figure out whether you're routing to an open-weight or a closed model — that determines whether you should be evaluating the provider on compute-price competition or on account-quota discount and stability. See our AI API relay comparison for specific providers.