If you need the state of this story in one line as of publication, it's: "announced, not live, no numbers." The "2x surge pricing during 9am–12pm and 2pm–6pm Beijing time" mechanism rumored in mid-July still shows up nowhere on DeepSeek's official pricing page as of this writing. Then, on August 6, DeepSeek posted a vaguer, broader notice through its user dashboard and API console — "we plan to raise overall pricing for DeepSeek API services in the near future, with a significant increase expected" — again with no specific numbers, no effective date, and no new rate card. That's DeepSeek's second pricing move in under a month, and it's the reason this piece exists: the window where the company has said "prices are going up" but hasn't said by how much is exactly when readers need a "what do I actually do now" guide — you don't need to wait for the final number to write something useful.
Contents
- 1. What we verified: peak-hour pricing still isn't live, and a vaguer new notice landed August 6
- 2. Re-checked before publishing: has anything changed as of August 7?
- 3. Why now: Liang Wenfeng's rationale and the sector-wide backdrop
- 4. How the community is reacting: mixed frustration in Chinese forums, alternative-shopping in English ones
- 5. Five technical mitigations developers are actually using
- 6. Do relay providers actually save you money on DeepSeek? Three mechanisms, three very different answers
- 7. What to actually do right now, if you're still on the fence
1. What we verified: peak-hour pricing still isn't live, and a vaguer new notice landed August 6
1.1 The most important question first: July's "peak-hour surge pricing" still isn't in effect
We pulled DeepSeek's official pricing page (api-docs.deepseek.com/quick_start/pricing/) directly to check: there's no trace of any peak-hour or time-of-day pricing anywhere on it. Both V4-Flash and V4-Pro are still flat-rate — V4-Flash's input runs $0.0028 on a cache hit and $0.14 on a cache miss, with output at $0.28; V4-Pro's input is roughly $0.0036 on a cache hit and $0.435 on a cache miss, with output at $0.87. Multiple independent sources (KuCoin, Dataconomy, and others) also stated around August 6 that "the official pricing page is still a single flat rate; peak-hour pricing hasn't gone live." In other words, the "2x surcharge during 9am–12pm and 2pm–6pm Beijing time" scheme announced back in mid-July remains "announced, not in effect" as of this writing. If you're using the DeepSeek API right now, you're still being billed the off-peak rate — your bill hasn't doubled because of this rumored mechanism.
1.2 But something new happened on August 6: a bigger, vaguer price-increase notice
On August 6, 2026, DeepSeek posted a notice through its user dashboard and API console (quoted verbatim by multiple outlets): "We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. The specific plan and effective date will be subject to official notice." This notice contains no specific numbers, no effective date, and no per-model breakdown — and its scope is broader than July's peak-hour scheme. Several reports point out that this implies even the off-peak base price could rise, not just "prices double during peak hours." Taken together, these are DeepSeek's two pricing moves within a single month, and the nature of the move has escalated from "time-of-day surcharge" to "a general repricing."
One limitation is worth stating plainly: as of this research and writing, no official or credible source has given a specific percentage increase, an effective date, or a new rate card.
We don't credit the "2x to 10x" figure circulating online
During research, we found an English-language SEO blog claiming DeepSeek's increase would be "2x to 10x," citing remarks attributed to "DeepSeek founder Jun Song." DeepSeek's actual founder is Liang Wenfeng — that piece gets the founder's name wrong, which is a clear signal of an unreliable source. We don't adopt its specific multiplier; we mention it only as a counterexample of the exaggerated speculation circulating online, and readers should be careful about source quality. This number does not appear anywhere in this article's conclusions as fact.
2. Re-checked before publishing: has anything changed as of August 7?
Pre-publication re-check: the official pricing page still hasn't been updated, and "a significant increase is expected" is now written into the pricing documentation itself
Because this is a live, fast-moving story, we ran an independent check right before publishing, re-pulling the official pricing documentation directly. Result: as of August 7, both V4-Flash and V4-Pro still show the same flat rates listed above, with no trace of peak-hour pricing anywhere. We also noticed a detail our earlier research hadn't caught: the line "we plan to raise overall API pricing in the near future, with a significant increase expected" is now written directly into the official pricing documentation itself — not just scattered across dashboard pop-ups and press paraphrases — which suggests this isn't a one-off announcement but DeepSeek's current, standing public position. Several independent reports published after August 6 (including Bloomberg and the South China Morning Post) cross-check to the same conclusion: DeepSeek has confirmed a broad increase is coming, but none of them have a specific number or effective date either. In other words, nothing substantive has changed between when we finished the underlying research and when this piece went live — it's still "confirmed to be coming, amount and timing both unknown." Treat this article as a snapshot of an information vacuum, not a settled conclusion that only becomes useful once official numbers land.
3. Why now: Liang Wenfeng's rationale and the sector-wide backdrop
More worth recording than "how much" is the pricing logic DeepSeek's own leadership has offered. This comes from founder Liang Wenfeng's remarks at a July investor briefing, quoted by multiple outlets including The Star Malaysia: DeepSeek's API pricing is based on a "reasonable profit model" — after buying a batch of hardware, the company wants to recoup that investment in roughly 10 months; Liang stated plainly that demand at current price levels is "almost inelastic" — "even a 50% price increase wouldn't meaningfully change token consumption"; and the company isn't chasing profit maximization, just a "reasonable return." In other words, this increase reads less like a company running out of money and more like a deliberate repricing decision made after estimating demand elasticity.
The backdrop is worth putting in the same frame: multiple reports say DeepSeek-V4-Flash's daily token consumption hit 8 trillion on August 1, widely cited as the immediate trigger that tipped the "low prices aren't sustainable" calculus. DeepSeek is also said to be closing a $740 million second funding round and preparing for an IPO within the year, while planning roughly 1GW of data-center capacity in Inner Mongolia — infrastructure spending on the order of $50 billion. And crucially, this isn't unique to DeepSeek: Zhipu's GLM raised prices 83% in Q1 2026 (even as usage grew 400% in the same period), and Moonshot's Kimi K3 paused new-user signups in July due to a demand surge. The "cheap prices for scale" playbook is retreating across China's large-model sector as a whole, in favor of a "sustainable profitability" narrative — this increase should be read against that sector-wide backdrop, not treated as an isolated event.
The only price change confirmed to be actually in effect right now is zero — the official rate card looks the same as it did a week ago. What has genuinely changed is expectations: the market now knows a price increase is coming and could be substantial, but nobody — including DeepSeek itself — has put a number on it yet. That information vacuum is itself the most accurate description of where things stand, and it's the direct backdrop for the community reaction and mitigation strategies covered below.
4. How the community is reacting: mixed frustration in Chinese forums, alternative-shopping in English ones
4.1 Chinese-language communities (V2EX, Zhihu)
We found at least three relevant threads on V2EX and two widely-read Zhihu pieces on this. What we could pull showed a few consistent patterns: the mood is a mix of disappointment and resignation rather than collective outrage — one user joked "I just started using it yesterday, and today comes the price-hike notice"; others questioned whether this signals cash-flow pressure. There's also skepticism directed at DeepSeek's own communication — one user posted a claim that "the company denied a price hike," which was immediately countered with "the notice is sitting right on the API console homepage," touching off a debate about source credibility that itself suggests DeepSeek's own communications haven't been transparent enough, deepening community distrust. Some users compared DeepSeek to overseas models like Claude, but more as a lament that "it used to be absurdly cheap" than an announcement of an actual switch. A chunk of the griping targets DeepSeek's pricing history as a pattern — from "a permanent 75% cut" to "peak-hour surge pricing" to now "an across-the-board increase," which some users summarized as a familiar playbook: build a reputation on rock-bottom prices, then quietly claw the margin back. Notably, none of the threads we could pull contained concrete discussion of switching to relay providers or self-hosting — V2EX read more as an emotional and trust-level conversation than a solutions-level one, a contrast worth noting against the more technical discussion in English-language communities and dev blogs below.
4.2 English-language communities (Reddit, Hacker News)
On Reddit's r/DeepSeek, the most-cited alternative in price-hike discussion threads is GPT-5.6 Luna, partly because of its vision support — some users say they've already "pre-emptively switched" to it. Hacker News also has a dedicated discussion thread, and the title alone signals that international developers are paying attention to this — but due to rate limiting we weren't able to pull the actual comment content this round, and we're stating that limitation plainly rather than inventing comment details.
5. Five technical mitigations developers are actually using
More useful than venting is what developers are actually doing about it. Pulling from CSDN developer blogs, ofox.ai, and similar sources, the mitigation strategies that keep coming up fall into five categories:
- Usage audits and technical token-saving: reviewing historical token consumption, trimming system prompts, using RAG to narrow the context window, setting a
max_tokensceiling, and switching to streaming responses; - Tiered model routing by task complexity: routing simple Q&A, text polishing, and basic code completion down to V4-Flash, reserving V4-Pro for genuinely complex reasoning and code generation — instead of defaulting everything to the higher tier;
- Decoupling via a multi-model/multi-provider abstraction layer: not hardcoding DeepSeek as the default backend, building a layer that can dynamically switch providers and route to whatever's cheapest at the moment — using lightweight open models like Qwen2.5-Coder as a pre-filter is one concrete approach mentioned;
- Avoiding long-term lock-in: not signing any long-term contract or large prepayment locked to current pricing until DeepSeek publishes an actual new rate card, and isolating DeepSeek spend in its own line on a cost dashboard so that once an increase lands, its real impact is visible immediately;
- Evaluating self-hosting the open-weight model: if your call volume is high enough, evaluating whether to rent GPUs and self-host DeepSeek's open weights (via vLLM, TGI, or similar inference frameworks), converting "pay-per-token to DeepSeek" into "pay-per-GPU-hour to a cloud provider." This path is only viable because DeepSeek open-sourced the weights in the first place — which is exactly the mechanism the next section digs into.
Separately, an ofox.ai piece titled "DeepSeek V4 Flash: 6 Ways to Pay Less" offers six tips that are entirely about optimizing usage within the official rate card — improving cache-hit rates, checking cache-read price gaps when picking a third-party host, explicitly pinning the latest model version to avoid silently defaulting to an older one, watching for hidden quality loss from cheap hosts using low-precision quantization or a shrunken context window, turning off thinking mode when it's not needed to save output tokens, and picking an agent framework that minimizes token spend from replayed context. None of this involves arbitrage or relay-provider discounts — it's a different question from "how do I find a channel below the official price," which is what the next section actually addresses.
6. Do relay providers actually save you money on DeepSeek? Three mechanisms, three very different answers
This is the part our readers care about most. Cross-referencing this site's own reviews of 120-plus relay providers that mention DeepSeek, we specifically stress-tested the question "does routing DeepSeek through a relay actually save money?" The honest answer: yes, but it depends entirely on which kind of relay, and the results vary enormously — it's not accurate to say relays are generally cheaper than official pricing.
6.1 Mechanism one: open-weight self-hosted compute competition — real, but not cheaper on every dimension
DeepSeek open-sourced the V4 line's weights, so any third party with enough compute can self-host and resell it — a savings path that's specific to DeepSeek, since closed models like GPT or Claude have no equivalent. Our own testing shows SiliconFlow pricing DeepSeek-V4-Flash at roughly ¥0.15/¥0.15 (input/output), against DeepSeek's official uncached rate of roughly ¥1/¥2 — about 85% cheaper on input and 92.5% cheaper on output, a substantial gap. But V4-Pro isn't uniformly cheaper: SiliconFlow prices it around ¥2/¥8, versus an official uncached rate of roughly ¥3/¥6 — about 33% cheaper on input, but actually about 33% more expensive on output. That's a counterexample worth stating honestly: open-weight third-party hosting doesn't guarantee every model, on every billing dimension, is cheaper than official — you need to run the math against your own input/output token ratio rather than assume "open-weight hosting is always cheaper."
EasyRouter (endorsed by Cheetah Mobile CEO Fu Sheng) prices everything at 85% of official across the board, but DeepSeek V4 Pro specifically runs as low as 25% of the official price — the single largest, most credibly-backed (not an anonymous small site) discount we found for any one DeepSeek model. Separately, third-party price-comparison site cctest.ai's live comparison page lists several smaller relays with fewer reviews quoting even lower DeepSeek prices, generally 70–93% below official. Those figures deserve real skepticism: the industry's common "markup multiplier" framework holds that relays legitimately buying official quota in bulk and reselling it typically land at 0.8x–1.5x the official price; multipliers as low as 0.05x–0.3x (5%–30% of official) usually indicate shared keys, reverse-engineered access, or grey-market accounts — a gap that unlikely to come from a normal bulk-purchase discount. We don't endorse those listings here; we're reporting the comparison data and the risk alongside it.
6.2 Mechanism two: FX/quota arbitrage — unrelated to DeepSeek's own pricing, a generic trick
Our own review of TokenRiver gives a concrete example: it settles at "1¥ = $1," while the actual market FX rate is roughly ¥7.2/$ — which works out to roughly 86% "savings" on DeepSeek V4 Flash by that math. But it needs to be said plainly: this "savings" has nothing to do with DeepSeek's own pricing mechanics — the exact same 86% figure holds for Claude Opus 5 and GPT-5.6 on the same platform, because it's just "$X/M price × 1¥=$1" arithmetic. It's a generic FX-peg accounting trick that applies to any dollar-priced model, and it amounts to "your top-up buys more nominal dollar credit at the official conversion rate," not "this specific model is genuinely cheaper on this specific relay." Several other relays we've reviewed on this site, running FX rates around ¥1–2.5/USD, are other instances of the same mechanism.
6.3 What DeepSeek's own review page says: relays reselling official quota typically mark up, not discount
Our own DeepSeek provider review on this site includes a line worth quoting directly: "Many relay providers support DeepSeek models... pricing is transparent: the official rate is the baseline, and relays add a markup on top of it." That lines up exactly with the two mechanisms above: relays that genuinely beat the official price are almost always ones that self-host open weights (mechanism one); relays that simply "buy official DeepSeek account quota in bulk and resell it" are, in this site's own review, described as marking prices up, not discounting them. Same company, two paths: self-host and you can genuinely save; resell official quota and, by DeepSeek's own reviewers' account, you typically pay more — that contrast is the most honest thing we can tell readers here.
"Bulk pre-purchase to lock in pre-hike pricing": no concrete case found this round
We specifically checked whether relay providers offer a mechanism to bulk pre-purchase official quota to lock in pricing before an increase takes effect, and found no specific, citable relay-provider case proving this mechanism exists — what search results turn up reads more like common-sense speculation than a conclusion backed by real sourcing. To be clear: not finding evidence doesn't mean this mechanism definitely doesn't exist, just that we found no reliable public information supporting it. This is different from ordinary bulk-purchase discounting (buy more, pay less per unit), which is a standard, long-standing procurement pattern — "hoarding pre-hike pricing ahead of an increase" is a one-off, narrow-window arbitrage play, and it's worth noting that since DeepSeek hasn't even published an effective date for this increase yet, "hoarding the old price" currently has no concrete window to act on in the first place — no effective date means no "ahead of the deadline" moment to exploit.
7. What to actually do right now, if you're still on the fence
Put all of the above together, and the honest advice right now is: don't panic over a number that hasn't been published yet, and don't assume the blanket idea that "relays are cheaper" will automatically save you money. What's genuinely worth doing today is implementing the token-saving tactics covered above — cache-hit optimization, tiered task routing, provider-abstraction decoupling — first, since those pay off regardless of whether prices rise. At the same time, isolate DeepSeek spend as its own line item on a cost dashboard, so that once an official number lands, you can calculate its real impact immediately instead of scrambling through invoices after the fact.
Put together, this means
DeepSeek's price increase is still at the "confirmed coming, amount unknown" stage. As of this publication, any claim of a specific percentage increase lacks official backing and should be treated with caution.
- Developers continuing on DeepSeek's official API short-term → prioritize cache-hit optimization and tiered task routing, both of which save real money regardless of whether the increase lands, and avoid signing any long-term price-locked contract during this window.
- Developers considering a relay provider → first work out whether it self-hosts open weights or resells official quota. The former (cases like SiliconFlow and EasyRouter above) delivers genuine discounts on specific models and billing dimensions; the latter is typically a markup, by DeepSeek's own reviewers' account; FX-arbitrage "savings" have nothing to do with whether DeepSeek itself is raising prices.
- Teams with high volume and real infrastructure capacity → it's worth seriously evaluating self-hosting the open-weight model, converting per-token spend into per-GPU-hour spend, but only after mapping out your real token consumption profile first.
- Everyone → treat this article as an information snapshot as of around August 7, 2026. Once DeepSeek publishes actual numbers and an effective date, check the official pricing page directly — we'll follow up with a dedicated recap once the real figures land.