August 13, 2026, may go down as the quietest information-dense night in this round of the AI price war. Just after midnight, DeepSeek held no event and issued no press release — it simply updated its API documentation to DeepSeek-V4-Pro-0813, shipping the GA of its flagship exactly 111 days after the Preview. The same night, xAI — folded into SpaceX and now branding itself SpaceXAI — posted a single line: "Introducing Grok 4.6." One team edited a doc, the other posted a tweet, but both were betting on the same lane: selling a credible flagship at half the price — or less — of the closed Western premium tier. This article puts those two "half-price nukes" on the same table, lays out their real differences, and answers the question our readers actually care about: can you reach either through a mainland China relay today, and which should you pick?

A note on fact discipline before we start. DeepSeek-side specs, benchmarks, and pricing come from DeepSeek's official API changelog; Grok 4.6-side facts come from xAI's official announcement, Artificial Analysis' independent evaluation, and third-party reporting such as FoneArena. Where a figure is vendor-reported and not yet independently verified, we say so explicitly; where sources contradict each other and cannot be cross-checked, we say "unverifiable" rather than invent a number to make a table look more complete. The gap between what each vendor reports and what third parties measure is one of the main things this piece exists to unpack.

1. Same-night collision: one edited a doc, one posted a tweet

Let's get the timeline straight, because it says a lot by itself. DeepSeek launched the V4 family on April 24, 2026, with both Pro and Flash marked as "Preview." On July 31, DeepSeek first matured the smaller, faster-iterating V4 Flash into the 0731 public beta — the release we covered in our DeepSeek V4 Flash deep-dive (the same architecture, one fresh post-training pass, DeepSWE from 7.3 to 54.4) — while V4 Pro stayed in Preview. In the early hours of August 13, the flagship GA, DeepSeek-V4-Pro-0813, finally landed.

Almost the same hour, xAI shipped Grok 4.6, just over a month after Grok 4.5 (July 8). It's the first model released under the SpaceXAI banner after xAI was merged into SpaceX. Both companies deliberately skipped the press-event route, but for different reasons: DeepSeek has always preferred to let documentation speak; xAI put its energy into being immediately live inside its own ecosystem (Cursor, Grok Build, X). Two flagships updating the same night effectively put the competition theme of H2 2026 on the table: not "who has more parameters," but "who post-trains better and who prices harder per unit."

There's also a piece of context our readers should hold onto. DeepSeek's GA documentation simultaneously warned of a "significant" overall API price increase in the near future, and this site has already verified that the V4 Pro "peak-hour dynamic surge" rumored since mid-July has still not actually gone into effect as of publication. That means the pricing you're about to read is the current official price — but quite possibly the cheapest window for a while. Both facts (the hike warning and the non-live surge pricing) have no concrete numbers yet; we report them as facts and don't speculate.

Why both companies chose mid-August is worth a beat. The whole industry is mid-escalation: OpenAI cut prices on its lower tiers in late July, Anthropic and Google are trading releases of their own flagship tiers at $5-10/M input, and Chinese open-weight labs keep pushing the "what does intelligence have to cost" question lower. Against that backdrop, a same-night double launch of two "half-price flagships" reads less like coincidence and more like both camps planting their stake in the same price band on purpose. It also echoes something this site has been tracking all month: the competitive variable that matters in H2 2026 is unit economics — post-training quality and price per useful token — not parameter bragging rights.

2. Specs: 1.6T/49B vs ~1.5T — neither one is about parameter count

Hard specs first, then plain English.

SpecDeepSeek V4 Pro 0813Grok 4.6
ArchitectureMoE, ~1.6T total / ~49B activeMoE, ~1.5T total (per Musk; not fully independently verified)
Context window1M tokens500K tokens
Max output384K tokensNo stated output cap
ModalityText-primary; GA adds native image reasoning (DeepThink engine)Text + image input, text output
Reasoning levelsNon-Think / Think High / Think MaxLow / Medium / High (default) / XHigh
Concurrency500Not published

The two share a point worth calling out separately: Grok 4.6 did not scale up parameters. Per xAI's official description, 4.6 keeps the same ~1.5T-parameter V9 base as 4.5; nearly all the gains come from post-training — rebuilding SFT trajectories across domains with model-generated reasoning and engineering data, then running reinforcement learning across agent environments (coding, kernel optimization, web development, CAD). This is almost exactly the same playbook as the DeepSeek V4 family: keep architecture and parameter count fixed, and squeeze out capability by redoing post-training. Two companies, one night, one method — the most direct evidence yet that the center of gravity in H2 2026 has moved from "stacking parameters" to "tuning post-training."

On context, DeepSeek V4 Pro's 1M tokens clearly beat Grok 4.6's 500K — a real difference for long documents and whole-repo code refactors. The trade-off: DeepSeek's concurrency cap is only 500, so file a ticket before load-testing at scale. On modality, the GA of V4 Pro adds the image reasoning that the Preview lacked (third-party hands-on coverage flagged image handling as the Preview's most obvious weak spot), while Grok 4.6 has been natively multimodal since 4.5. That matters for the selection advice later: if your job needs to "read" a screenshot or design file, both can now do it, but Grok 4.6's multimodal support is more mature.

3. Benchmarks: why DeepSeek's official numbers look stronger, yet third-party AAII is 53 vs 61

This is the section to read carefully, because it contains the easiest trap of the whole piece: DeepSeek's self-reported agent scores jumped impressively, but Artificial Analysis gives V4 Pro an intelligence index (AAII) of just 53, versus 61 for Grok 4.6. The two sets of numbers aren't in conflict; they just mean "vendor-reported" and "third-party, one consistent methodology" are different things. We show both, with sources labeled.

First, DeepSeek's official 0813 changelog, comparing against the April Preview (all vendor-reported, no large-scale independent replication yet):

BenchmarkV4-Pro-Preview (Apr)V4-Pro-0813 (Aug)
DeepSWE12.862.7
Terminal-Bench 2.172.187.9
CyberGym52.783.3
NL2Repo38.561.5
DSBench-Hard31.167.2
DSBench-FullStack41.871.1
Toolathlon-Verified55.974.1
HLE (with tools)48.260.0
AutomationBench12.831.8
Agents' Last Exam16.525.7

Watch the DeepSWE row: 12.8 to 62.7, roughly a 5x jump — the same pattern as V4 Flash's "one retraining pass, 7x" story. DeepSeek's changelog also includes some rival comparisons: on DeepSWE, V4 Pro's 62.7 beats Claude Opus 4.8's 58.0; on Terminal-Bench 2.1, 87.9 beats Opus 4.8's 85.0; CyberGym's 83.3 edges Claude Fable 5's 83.1. But against the stronger Claude Fable 5, it still trails on Toolathlon (74.1 vs 77.9), DSBench-FullStack (71.1 vs 77.2), and DeepSWE (62.7 vs 70.0). All of these are DeepSeek's own numbers.

Now Grok 4.6. The independent evaluator Artificial Analysis puts its AAII at 61, tying OpenAI's GPT-5.6 Sol Max (61), one to two points behind Claude Opus 5 Max (63) and Claude Fable 5 Max (62). For reference, Grok 4.5 scored 56 and Grok 4.3 scored 38 — so 4.6 moved the third-party composite from 56 to 61 in about a month. Where 4.6 leads is agentic knowledge work: GDPVal-AA v2 at Elo 1753 is the highest among the public comparisons (Fable 5 Max 1741, GPT-5.6 Sol 1728); AA-Briefcase (long-horizon knowledge work) at Elo 1577 is Fable 5-tier; in the legal vertical, Harvey LAB at 15.8% blows past GPT-5.6 Sol's 2.5%. It also has clear weaknesses: DeepSWE v1.1 at 65.9% (a big gain from 4.5's 54%, but below Sol Max's 73% and Fable 5 Max's 70%), and Terminal-Bench v3.0 at 26%, well behind the ~34-35% of competitors.

Here we have to draw an honest line for readers: DeepSeek's "Terminal-Bench 2.1 87.9" and Grok 4.6's "Terminal-Bench v3.0 26" are not comparable. Different benchmark versions, different methodologies. Anyone who subtracts them and concludes "DeepSeek crushes Grok" is comparing apples to oranges. Likewise, DeepSeek's DeepSWE (62.7) and Grok 4.6's DeepSWE v1.1 (65.9) share a name but not a version or an evaluation protocol. That's precisely why the third-party AAII, run under one consistent methodology, is the cleaner reference: in Artificial Analysis' eyes, Grok 4.6 is 8 points ahead of DeepSeek V4 Pro's 53. As for which vendor's self-reported numbers are more trustworthy, we'll wait for independent replication — until both show up on public leaderboards like LMSYS Chatbot Arena, we treat both sets of self-reported scores as "reported but not yet verified."

How to read this benchmark collision sensibly

Vendor-reported agent scores jumping fast means post-training is genuinely improving, but it doesn't substitute for a third-party comparison under one methodology. The cleanest third-party data point right now: AAII 61 for Grok 4.6, AAII 53 for DeepSeek V4 Pro. If you only trust third parties, the conclusion is clear; if you want to know how high each vendor claims to run on its own agent benchmarks, DeepSeek's self-reported table is worth recording too. Keep the two sets separate.

4. Pricing: a ~7x output gap and a 100x+ cache-hit gap

If the benchmarks were a mixed bag, pricing is the part of this collision you can read at a glance. The prices below are official first-party API prices (not relay discounts); DeepSeek's are in RMB with a USD conversion, Grok's in USD:

ModelContextInput $/MCache hit $/MOutput $/M
DeepSeek V4 Pro 08131M$0.435 (¥3)≈$0.0036 (¥0.025)$0.87 (¥6)
DeepSeek V4 Flash 07311M$0.14≈$0.0028$0.28
Grok 4.6 (prompt <200K)500K$2.00$0.50$6.00
Grok 4.6 (prompt ≥200K, doubled)500K$4.00$1.00$12.00
GPT-5.6 Sol (reference)$5.00$30.00
Claude Opus 5 (reference)$5.00$25.00
Claude Fable 5 (reference)$10.00$50.00

Three conclusions fall out of this table. First, DeepSeek V4 Pro is roughly 1/7 the output price of Grok 4.6 ($0.87 vs $6.00) and about 1/4.6 the input price ($0.435 vs $2.00); against Claude Fable 5 ($10/$50), V4 Pro's output is about 1/57, which is where the "about 2% of Fable 5's price" line comes from. Second, the cache-hit gap is even more extreme: DeepSeek V4 Pro's cache hit is ~$0.0036/M, Grok 4.6's is $0.50/M — more than 100x apart. If your application has heavily reused prefixes and leans on prompt caching, that single line decides your bill.

Grok 4.6's pricing isn't random either. Holding at $2/$6, it positions itself as "about half the price of comparable frontier models" — relative to GPT-5.6 Sol ($5/$30) and Claude Opus 5 ($5/$25), that's 60%+ cheaper on output. So the two "half-price" labels reference different baselines: DeepSeek is cheap in absolute terms against every closed flagship, while Grok 4.6 is half the price of its same-tier Western peers. One more detail easy to miss: prompts over 200K tokens double Grok 4.6's input and output ($4/$12), and its cache-hit price rose from $0.30 in 4.5 to $0.50 — so long-context, cache-heavy workloads will cost more than the headline numbers suggest. Model your own prompt length before committing.

And one more concrete calculation to make the "7x" tangible. Suppose a deep-analysis job over a 100K-token document that outputs ~20K tokens: with DeepSeek V4 Pro, input ≈ $0.435 × 0.1 ≈ $0.044 and output ≈ $0.87 × 0.02 ≈ $0.017, so under $0.07 total; with Grok 4.6, $2 × 0.1 = $0.20 input and $6 × 0.02 = $0.12 output, so $0.32 — roughly 4.6x. That doesn't even count cache hits, where DeepSeek's hit price is nearly negligible while Grok 4.6 accrues $0.50/M on every reused prefix. To be fair, Grok 4.6 does have a cost-per-task edge of its own: Artificial Analysis estimates ~$0.84 per long-horizon AA-Briefcase task, thanks to extreme token efficiency (~53 turns and ~0.5B input tokens, versus ~103 turns and ~2.0B for Claude Opus 5 Max), placing it on the intelligence-vs-cost Pareto frontier. For budget-sensitive domestic apps, though, DeepSeek's absolute low price usually wins.

On the DeepSeek side, one unknown has to be restated: the official documentation warns of a "significant" overall API price increase soon, and the rumored "peak-hour dynamic surge" still isn't live as of publication. In other words, the ¥3/¥6 prices written here are the current official price but probably not the long-term one — the same caveat we've hammered on in our DeepSeek price-hike tracker, so we won't repeat it at length here.

5. Ecosystem: OpenAI + Anthropic compatibility vs the SpaceXAI family

Beyond specs and price, the part of this collision that best shows the two are fighting different wars is ecosystem strategy.

DeepSeek V4 Pro plays the "open ecosystem, dual-protocol compatibility" card. The official API speaks OpenAI's Chat Completions and Responses API, plus Anthropic's Messages format — no adapter layer needed to point Claude Code at DeepSeek via an environment variable, or to plug in tools like Codex that natively use the Responses API. DeepSeek's weights are MIT-licensed on Hugging Face, so any relay provider or self-hosted cluster can deploy them. For domestic developers, the migration cost is close to zero: change a base_url and a key, done.

Grok 4.6 plays the "own ecosystem, live-day availability" card. It became the default model in Cursor and Grok Build on day one, and went live on xAI's API, OpenRouter, Vercel, and Cloudflare with doubled credits for the first week. It rides Musk's personal megaphone in the consumer-facing X/Grok experience and is a "just works" option inside tools like Cursor. The cost: it's closed-weight, can't be reached from mainland China without a relay, and its strengths live mostly in the Western ecosystem and consumer-side use cases.

Read through the lens of this site's audience — domestic developers and relay users — and the picture is clear: DeepSeek's "dual-ecosystem compatibility + direct China access + MIT open weights" is almost tailored for you; Grok 4.6 requires a relay or OpenRouter for stable access, and its appeal is concentrated in "you want the higher third-party-scored flagship and you're OK with going through a relay layer." That's the foundation for the verdict below.

6. Relay-provider reality: who's already serving them, who hasn't updated yet

DeepSeek V4 Pro first: the most direct entry is always the official platform, platform.deepseek.com — directly reachable from mainland China with no proxy. The model name is still deepseek-v4-pro, so existing code doesn't change; DeepSeek swapped the weights server-side. Among the providers this site has verified, plenty of listed relays already carried deepseek-v4-pro, including SiliconFlow (硅基流动), Laozhang API (老张API), YKH.AI, FlowBar, UiUiAPI, plus OpenRouter, CloseAI, 4SAPI, 302AI, AIHubMix, RunAPI, DuckCoding, and TokenRiver. Because this upgrade didn't change the model ID, any relay that forwards the official API should be serving the 0813 build automatically; if a relay self-hosts DeepSeek's open weights instead, it has to follow the new weights on its own — check each vendor's announcement.

Grok 4.6 next. Officially it's confirmed live on the xAI API, Cursor, Grok Build, and OpenRouter. For domestic relay readers, the pragmatic check is whether your relay/gateway has actually added the grok-4.6 model ID. Grok coverage among this site's reviewed vendors has never been scarce — OpenRouter, CloseAI, 4SAPI, 147API, Shiyun API (诗云API), AIHubMix, RunAPI, APIYi, PoloAPI, DuckCoding, FlowBar, 302AI, TokenRiver, and Shenma API (神马API) have all listed Grok 4.3/4.5 at various points, and RelayDance even specializes in Grok relay access ("few platforms in China can reach Grok stably"). But "has carried Grok 4.5" is not "has updated to Grok 4.6" — 4.6 is a day old, so you'll need to check each console for the grok-4.6 ID and its price before routing real traffic.

If you use a gateway like LiteLLM or OpenRouter, the wiring is unchanged — just swap in the right model IDs:

model_list:
  - model_name: deepseek-v4-pro
    litellm_params:
      model: openai/deepseek-v4-pro
      api_base: https://api.deepseek.com
      api_key: os.environ/DEEPSEEK_API_KEY
  - model_name: grok-4.6
    litellm_params:
      model: grok/grok-4-6
      api_key: os.environ/XAI_API_KEY

A necessary caveat: those IDs are the official/mainstream-gateway conventions, not a guarantee that every relay has added grok-4.6. DeepSeek's side is low-risk because the model ID didn't change; Grok's side is brand new, so go by what each relay's console actually lists before pointing production traffic at it.

7. The verdict: which to pick for which job

  • Domestic developers, budget-sensitive, need 1M context / long docs / whole-repo code → DeepSeek V4 Pro, directly. Direct China access, dual-ecosystem compatibility, 1/4.6 to 1/7 of Grok's price; just remember the "price hike coming soon" warning — if you want to lock in pricing, do it early.
  • Batch apps that lean heavily on prompt caching → DeepSeek V4 Pro, no contest: its cache-hit price (~$0.0036/M) is 100x+ cheaper than Grok 4.6's $0.50/M.
  • You want the highest third-party score among flagships, native multimodal, agentic knowledge work → Grok 4.6. AAII 61 ties GPT-5.6 Sol Max and it leads GDPVal-AA, but it costs 7x+ more and needs a relay/OpenRouter.
  • Interactive front-end generation, Cursor / Grok Build workflows → Grok 4.6 is the default model and works out of the box; this is its most natural home.
  • High-volume, low-complexity batch and sub-agent tasks → stick with V4 Flash ($0.14/$0.28, 2,500 concurrency) and reserve Pro for genuinely hard tasks; the two-tier combo is the cheapest.
  • You want self-hostable, vendor-portable, no ecosystem lock-in → DeepSeek (MIT open weights) is the only option; Grok 4.6 is closed.

8. Conclusion: the real signal of that night

DeepSeek V4 Pro GA and Grok 4.6 landing the same night looks like two "half-price nukes" colliding, but it actually reveals two converging trends for H2 2026.

Trend one: the center of gravity has moved from "stacking parameters" to "post-training plus unit price." Both companies explicitly said this generation doesn't change the architecture or parameter count — the gains came from redoing post-training — and both pinned their pricing at "half or less of the closed Western premium tier." While OpenAI and Anthropic sit at $5-10/M input and $25-50/M output, DeepSeek and Grok are delivering flagship-class ability at $0.44-2 input. For budget-sensitive applications, this price band just got crowded — and, for the first time, genuinely comparable.

Trend two: what looks like a collision is mostly two positions in the same battle. DeepSeek is the price/performance king on "absolute low price + open ecosystem + direct China access," playing for cost and ecosystem freedom; Grok 4.6 is the "half the price of Western peers + SpaceXAI family + highest third-party score" play on the Western/consumer side, playing for experience and integration. Their real intersection is that both force OpenAI, Anthropic, and Google's flagship pricing to be re-examined — and the choice between them comes down to where your workload sits on the "cost, openness, direct access" vs "score, experience, ecosystem" axis.

Putting it together, that means

  • Domestic + budget-sensitive + long-context / dual-ecosystem → DeepSeek V4 Pro; lock in the price before the hike, direct connection is enough.
  • You want the highest third-party score + native multimodal + Cursor/Grok Build → Grok 4.6, via a relay or OpenRouter; watch the doubling above 200K prompts.
  • Pure batch / sub-agent workloads → still route to V4 Flash for the best cost.
  • Two reminders: DeepSeek has already flagged a near-term overall price hike and its surge pricing isn't live; Grok 4.6 is a day old, so verify each relay's grok-4.6 model ID before putting production traffic through it.