On August 22, 2026, BlockBeats reported that DeepSeek had released V4-Flash-Vision-Exp, a vision-enabled version of its V4-Flash model, alongside its first Agent benchmark scores — a 2:2 tie with Anthropic's flagship Claude Opus 4.8 across four multimodal Agent evals. This is not a new flagship. It is "eyes" bolted onto the workhorse. For the mainland China relay ecosystem, the most consequential detail is price: if the vision variant inherits V4 Flash's peak/off-peak tier (off-peak ¥1.5/¥4.5, peak ¥3/¥9 per M tokens), it becomes one of the cheapest "screenshot-and-UI-literate" multimodal Agent models on the market. This article breaks the variant down from five angles — positioning, specs, benchmarks, pricing, and relay availability — and is explicit about which numbers are verified and which are not.
First, the factual discipline and time base. This article was fact-checked as of August 22, 2026. The benchmark numbers are relayed from BlockBeats' summary of DeepSeek's own release materials and are DeepSeek's self-reported figures, not an independent third-party leaderboard — this article treats them as vendor claims and labels them accordingly. The base V4 Flash specs (284B/13B activated MoE, DSA hybrid sparse attention, 1M context) and the peak/off-peak tier follow data this site already verified and published (deepseek.md vendor page, lastVerified 2026-08-15; official tiering effective 2026-08-16 16:00 UTC). The vision variant's exact API model ID, and whether vision input is billed separately, were not verifiable in official documentation at publication time — wherever a number cannot be cross-checked, this article says so plainly instead of inventing data to make a table look complete.
Contents
- 1. Positioning: Not a New Model, a Multimodal Extension of V4 Flash
- 2. Specs & Capabilities: 284B/13B MoE + 1M Context, Plus "Eyes"
- 3. Benchmarks: First Agent Scores, 2:2 vs Opus 4.8 — How to Read Them
- 4. Pricing: Will Vision Input Be Billed Separately? Cost Under Peak/Off-Peak
- 5. Relay Availability: Who's Selling deepseek-v4-flash-vision
- 6. Price Comparison: Base V4 Flash Relay Rates as a Reference
- 7. How to Access: OpenAI-Compatible Code and Model-ID Caveats
- 8. Selection Verdict: Who Should Wait for It, Who Shouldn't
1. Positioning: Not a New Model, a Multimodal Extension of V4 Flash
First, place V4-Flash-Vision-Exp inside DeepSeek's lineup, because that sets the coordinate system for everything that follows. DeepSeek's V4 series currently has two tiers: V4 Flash (the volume workhorse, 284B/13B activated MoE) and V4 Pro (the flagship, 1.6T/49B activated MoE). Since V4 Flash graduated from Preview to the 0731 public-beta build on July 31, it has been a pure-text model — this site's late-July review stated it plainly: it cannot look at images, read screenshots, or reconstruct UIs, and if you needed multimodality back then, the recommendation was Qwen3.6-35B-A3B. V4-Flash-Vision-Exp exists to close exactly that gap.
Each part of the name matters. "V4-Flash" says the base is the Flash tier, not a new flagship; "Vision" says the addition is visual/image input; "Exp" (Experimental) says this is an experimental variant, not a stable release — in DeepSeek's pattern, experimental builds are offered to developers first and promoted based on feedback (the 0731 build took the same Preview→public-beta path). On the available evidence, the goal is "bring multimodal Agent ability onto V4 Flash's low-cost base," not "build a model stronger than Pro." This mirrors how Google treats its Flash tier — flagship for the brand, Flash for developer volume — and "can see images" is now drifting from a flagship perk to a workhorse baseline.
There is a signal worth noticing in the timing. DeepSeek had just put peak/off-peak pricing on the V4 series in mid-August (off-peak at half price), explicitly to nudge batch load into off-peak hours and smooth out compute pressure. Releasing a vision variant in the same month tells you it is not content to be merely "the cheapest text-only model" — it wants "cheap + multimodal + Agent" all in one price band. For relay readers, that means if you have been forced onto pricier multimodal models because V4 Flash could not see, the vision variant is a "cheap eyes" release genuinely worth waiting for.
2. Specs & Capabilities: 284B/13B MoE + 1M Context, Plus "Eyes"
First, the base specs (already verified in this site's coverage of the V4 Flash 0731 build): 284B total parameters / 13B activated in a MoE, DSA (hybrid sparse attention) architecture, roughly 1M token context, text-only in and out. DeepSeek has not published new parameter counts for V4-Flash-Vision-Exp. Given how DeepSeek treats experimental builds (the 0731 update famously changed "not a single byte of architecture" and simply re-ran post-training), the most likely reading is that the vision variant adds a vision encoder plus multimodal alignment training on the same 284B/13B base, rather than training a much larger model — but that is inference, not something DeepSeek has confirmed, and this article labels it as such.
The capability change is concentrated in the input modality: the vision variant adds the ability to understand images, screenshots, UI mockups, charts, and scanned documents. For Agent workflows this is a real upgrade — a text-only V4 Flash Agent could only read textual descriptions; once it hits "read this error screenshot," "match this UI mockup," or "extract the numbers from this chart," it was stuck. The vision variant brings those tasks into scope. To be clear, what DeepSeek announced is vision input; there is no mention of image generation or video understanding, and output remains text. It is "an Agent model that can see," not "a model that draws."
One spec question matters a lot to relay users: are context and output length the same as the base model? Because vision tokens consume context (a single image can cost hundreds to thousands of tokens), the effective text headroom under 1M context may shrink slightly when images are mixed in. DeepSeek has not published numbers on this, so this article will not guess — just note that you should test the "long document + images" ceiling once, before putting it into production.
3. Benchmarks: First Agent Scores, 2:2 vs Opus 4.8 — How to Read Them
Read this section carefully, because it is where the pitfalls live: all of the following scores are DeepSeek's self-reported numbers (relayed via BlockBeats), with no large-scale independent re-test as of publication. This article relays them as vendor claims, not independent facts.
The first Agent benchmark release compares the variant against Anthropic's flagship Claude Opus 4.8 across four multimodal Agent evals, ending 2:2:
| Multimodal Agent Benchmark | V4-Flash-Vision-Exp | Opus 4.8 | Winner |
|---|---|---|---|
| Agents' Last Exam | 27.3 | 25.7 | DeepSeek |
| ZeroBench | 35.0 | 34.0 | DeepSeek |
| ApexBench | 36.5 | 39.4 | Opus 4.8 |
| Chartography | 64.3 | 65.0 | Opus 4.8 |
In other words: DeepSeek edges ahead on Agents' Last Exam (27.3 vs 25.7) and ZeroBench (35.0 vs 34.0), and Anthropic edges ahead on ApexBench (36.5 vs 39.4) and Chartography (64.3 vs 65.0). Every margin is 1–3 points — less "a clean tie" than "indistinguishable within noise." That itself is the notable signal: a workhorse-tier 284B/13B-activated model, on the vendor's own numbers, is trading blows with a flagship. But keep repeating: this is self-reported, methodologies differ between vendors, and cross-vendor comparisons deserve caution.
The more instructive comparison is how much the variant improved over the text-only V4-Flash-0731:
| Benchmark | V4-Flash-0731 (text-only) | V4-Flash-Vision-Exp | Change |
|---|---|---|---|
| ApexBench | 26.2 | 36.5 | +10.3pp |
| Agents' Last Exam | 25.2 | 27.3 | +2.1pp |
| DeepSWE (text Agent) | 54.4 | 59.3 | +4.9pp |
Two takeaways. First, vision input is a qualitative jump for "look-at-the-screen / operate-a-UI" Agent tasks: ApexBench leaps from 26.2 to 36.5, a 10-point gain — exactly the kind of task where a text-only model could not even participate, so the headroom is largest. Second, text Agent ability was not sacrificed and even ticked up: DeepSeek says it leads the text-only build in 6 of 7 text evals, with DeepSWE rising from 54.4 to 59.3, even surpassing Opus 4.8's 58.0. The vision-alignment training did not cost coding ability; like the 0731 retrain, it quietly pushed text-Agent ability up a notch as a side effect.
How to think about these scores
The self-reported "2:2 with Opus 4.8" is credible as a direction — "the vision variant brings V4 Flash's multimodal Agent ability close to flagship level" — but it is not a substitute for a unified third-party measurement. ApexBench +10.3pp and DeepSWE 54.4→59.3 are strong evidence that "adding vision made it better"; "tying Opus 4.8" is vendor framing that needs to be re-checked once third-party boards (Artificial Analysis, LMSYS, etc.) update. Treat vendor-reported and third-party numbers as two parallel data sets and never mix them — that is the single most important message of this article.
4. Pricing: Will Vision Input Be Billed Separately? Cost Under Peak/Off-Peak
Pricing is the most practical part of the vision variant for relay readers. Start with the base V4 Flash official price currently in effect (DeepSeek moved the V4 series to a peak/off-peak two-tier system effective 2026-08-16 16:00 UTC, per million tokens):
| Tier | Input ¥/M | Output ¥/M | ≈ USD |
|---|---|---|---|
| Off-peak (≈half price) | ¥1.5 | ¥4.5 | $0.22 / $0.66 |
| Peak | ¥3 | ¥9 | $0.44 / $1.32 |
Peak hours are 01:00–04:00 and 06:00–10:00 UTC daily. For reference, before the tiering took effect, V4 Flash was $0.14/$0.28 per M tokens (input/output) — so the off-peak price is roughly 57% above the old input price, but half the peak price. For budget-sensitive workloads, off-peak scheduling remains the first cost lever.
Now the vision variant. Here honesty is required: as of publication, DeepSeek's official docs do not list a separate price sheet for V4-Flash-Vision-Exp, and whether vision input is billed differently — or whether it inherits the peak/off-peak tier — is unverified. Two reasonable inferences (inference, not verified fact): first, it will very likely reuse V4 Flash's peak/off-peak tier — experimental builds usually run inside the workhorse tier's pricing frame, and carving out a new price band just adds migration cost; second, vision input will most likely be converted into image tokens and folded into the input price, rather than billed per image — that is the mainstream convention in the OpenAI-compatible ecosystem. Both points need to be confirmed against DeepSeek's API docs; this article makes no assertion.
A quick cost model, assuming "same price as base." Suppose a "look at this UI mockup and change the front end" Agent task: input is one UI screenshot (~1,000 tokens after encoding) plus 20K tokens of code context; output is ~10K tokens. At off-peak ¥1.5/¥4.5: input ≈ ¥0.03, output ≈ ¥0.045, under ¥0.08 total for the task. Even at peak ¥3/¥9, it is ~¥0.16. No "can-see-images" closed multimodal model (typically $1/M input and $5–15/M output) can touch that cost curve. If the vision variant does land at this price, it would be a step-change for high-volume "screenshot understanding + Agent automation" workloads.
5. Relay Availability: Who's Selling deepseek-v4-flash-vision
For mainland China readers, the practical question is whether relays can serve the vision variant yet. Honest status: as of publication (2026-08-22), none of the 153 AI API providers cataloged and verified on this site explicitly lists deepseek-v4-flash-vision or any vision-variant model ID in its model list. That is entirely expected — the variant shipped the same day, and the relay pipeline from "official availability" to "listed in my console" typically takes hours to days. So you will almost certainly not find it in any relay console yet. What this article can give you is "who is most likely to list it first" and "what it will probably cost once listed."
Three kinds of relays are the likely early movers. First, providers that advertise fast new-model sync: ProAI API (proaiapi.tech, whose pricing notes explicitly say "new models sync fast: grok-4.6/gemini-3.7-flash/qwen3.8-max already live," lastVerified 2026-08-16) and SiliconFlow (claims DeepSeek models update "nearly in sync with official," lastVerified 2026-08-11). Second, large-catalog aggregators: NoDAPI (545+ models), JENIYA (454+ models), and MKEAI (all lastVerified 2026-08-16) — their DeepSeek coverage already spans V4 Flash/Pro, so adding a vision variant is one more line. Third, DeepSeek-focused discount relays: RunAPI, 4SAPI, 302AI, AIHubMix, etc., which tend to follow quickly once official availability opens. Fair warning: these are "most likely," not "already live" — always check each relay's actual console list and price, and run a small test before loading real traffic.
On the model ID: the base model uses deepseek-v4-flash (the legacy alias deepseek-chat was retired July 24). For the vision variant, this article can only confirm the product name V4-Flash-Vision-Exp from the report — the exact API model ID (possibly deepseek-v4-flash-vision, deepseek-v4-flash-vl, or a similar variant) was not verifiable in official docs at publication time. Relays usually copy the official ID verbatim, but some use custom aliases (e.g., deepseek-v4-flash-vision-exp or deepseek/deepseek-v4-flash-vision). Before production, confirm against the official API docs and each relay's model list — do not point real traffic at an ID this article merely suggests.
6. Price Comparison: Base V4 Flash Relay Rates as a Reference
Because no relay has priced the vision variant yet, the table below shows each provider's current published rate for the base deepseek-v4-flash (data verified when this site cataloged each provider, with the verification date noted) as the reference line for what the vision variant will likely cost. The relay industry typically applies the same per-token unit price to different variants of the same model (the vision variant just bills a few more input tokens for the image), so this base price is the closest available estimate.
| Provider | Input ¥/M | Output ¥/M | Notes (verified date) |
|---|---|---|---|
| DeepSeek official (off-peak/peak) | ¥1.5 / ¥3 | ¥4.5 / ¥9 | Peak/off-peak tier, effective 2026-08-16 (verified) |
| SiliconFlow | ¥1.00 | ¥2.00 | Cache ¥0.02; verified 2026-08-11, predates the tier, may have changed (unre-checked) |
| RunAPI | ≈ ¥0.77 | ≈ ¥1.54 | Verified 2026-08-15; FX ≈ ¥6.75–6.79/$1 (verified) |
| OpenRouter | ≈ $0.068 (≈¥0.46) | ≈ $0.168 (≈¥1.14) | Verified 2026-08-11; aggregate, price varies by routing (verified) |
| 4SAPI / 302AI / AIHubMix / FlowBar | — | — | List deepseek-v4-flash in catalog; per-model price in console (not published) |
Three conclusions. First, relay prices sit below the official off-peak price: SiliconFlow at ¥1.00 input is a third cheaper than the official off-peak ¥1.5, and RunAPI (¥0.77) and OpenRouter (≈¥0.46) are well below — consistent with this site's earlier cross-vendor findings that relays apply a "discount + FX" double spread on DeepSeek volume models. Second, even if the vision variant reuses this price, it stays far below any image-capable closed multimodal model: even at the highest official peak price (¥3/¥9), it is a fraction of Claude/GPT vision pricing. Third, the peak/off-peak tier matters even more for the vision variant: if vision input is billed as tokens, off-peak scheduling saves more on "seeing" workloads, because screenshots inflate the input share — and input is the tier with the most pronounced peak/off-peak gap.
Repeating the honest labels: the SiliconFlow and OpenRouter quotes were verified 2026-08-11, before the official tier took effect (08-16), and whether they have since adjusted or added the vision variant has not been re-checked; RunAPI was verified 08-15, also before the tier. This article treats any relay quote for the vision variant itself as "not yet available" and will update once relays list it. Whatever you are reading, the live console price wins.
7. How to Access: OpenAI-Compatible Code and Model-ID Caveats
Access is identical to every OpenAI-compatible relay this site has covered before — the only differences are the model ID and an extra image field in the input message. Using the OpenAI SDK (model ID shown as deepseek-v4-flash-vision as a placeholder; use the official ID):
from openai import OpenAI
client = OpenAI(
api_key="your-relay-api-key",
base_url="https://your-relay-domain/v1" # use the relay console's actual base_url
)
response = client.chat.completions.create(
model="deepseek-v4-flash-vision", # use the official model ID
messages=[
{"role": "system", "content": "You are a front-end engineer. Propose changes based on the screenshot."},
{
"role": "user",
"content": [
{"type": "text", "text": "Analyze this UI screenshot and suggest improvements."},
{"type": "image_url", "image_url": {"url": "https://example.com/ui.png"}}
]
}
],
)
print(response.choices[0].message.content) If you use LiteLLM or a gateway aggregator, the logic is the same — just point model at the vision variant:
model_list:
- model_name: deepseek-v4-flash-vision
litellm_params:
model: deepseek/deepseek-v4-flash-vision
api_key: os.environ/DEEPSEEK_API_KEY Important: deepseek-v4-flash-vision is a placeholder, not a verified official model ID. The variant shipped hours ago; the official API docs may not even have the model listed yet, and no relay can have it live that fast. By the time you read this, the correct flow is: confirm the model ID in DeepSeek's official API docs, then check each relay console, then run a small paid test to confirm image input actually works — the classic failure mode is a relay that proxies only the text channel and silently drops the image field. Before production, run one image-plus-text call and verify the reply is not a "text-only answer that ignored the image."
8. Selection Verdict: Who Should Wait for It, Who Shouldn't
- Screenshot/UI/chart Agent tasks + budget-sensitive → this is the sweet spot. If you are currently paying $1/M+ closed multimodal models just to "see" images, wait for the vision variant to hit relays; costs may drop by an order of magnitude. Test the image channel first.
- Pure-text batch workloads → do not wait. The base
deepseek-v4-flashis already enough, at the lowest price with the most mature ecosystem. - Need the strongest multimodal flagship intelligence → do not use an experimental build against a flagship. The variant ties Opus 4.8 only on vendor-reported numbers, and still trails on ApexBench (deep GUI operation); if you have hard accuracy requirements, wait for third-party boards to update.
- Three things to verify before production → ① the official and relay model IDs match; ② image input is genuinely processed, not silently ignored; ③ the real bill after peak/off-peak pricing and image-token conversion.
9. Conclusion: The Real Signal of the Vision Variant
DeepSeek V4-Flash-Vision-Exp is an "experimental" release, but the signal it sends is bigger than its experimental status. Three things are worth remembering.
First, "cheap multimodal Agents" are becoming real. The V4 Flash workhorse tier plus peak/off-peak pricing (off-peak ¥1.5/¥4.5) was already among the lowest prices on the market; now, with vision input, it is trading blows with Opus 4.8 (2:2) on the vendor's own numbers. If the vision variant inherits that price, the cost ceiling for "an Agent model that can read screenshots" drops by an order of magnitude in one move — a more structural change than any single benchmark. Second, the vision-to-Agent gain is real: ApexBench 26.2→36.5 shows that "seeing the screen" converts directly into "operating the interface," and text-Agent ability (DeepSWE 54.4→59.3) rose as a bonus. But third, every score is self-reported, and the variant is still experimental: the model ID is unsettled, separate vision billing is unverified, and no relay has listed it. By the time you read this, trust the official API docs and each relay console — confirm the ID, test the image channel, then scale.
A closing note for relay readers: the vision variant is the most worth-watching "cheap multimodal entry point" of August 2026. You do not need to wire it up today — you cannot, yet. But it belongs on your watchlist: once fast-syncing relays like SiliconFlow, ProAI API, or NoDAPI list deepseek-v4-flash-vision at base-model pricing, it becomes the most extreme cost-performance choice for screenshot understanding, UI Agents, and document-multimodal parsing at high call volume.
Putting it together
- What it is → a multimodal extension of V4 Flash (284B/13B MoE, 1M context), experimental, adding image input; output remains text.
- Benchmarks → self-reported; 2:2 with Opus 4.8 across four multimodal Agent evals; ApexBench 26.2→36.5, DeepSWE 54.4→59.3. Independent re-tests still pending.
- Pricing → no separate official price sheet yet; most likely inherits V4 Flash's peak/off-peak tier (off-peak ¥1.5/¥4.5, peak ¥3/¥9); separate vision billing unverified.
- Relays → none of the 153 cataloged providers lists the vision variant yet; fast-syncing relays (SiliconFlow, ProAI API, NoDAPI, JENIYA) are the likely first movers; verify the image channel before production.