On August 14, 2026, Google officially released Gemini 3.7 Flash. This is not another slideware model — it went live the same day in Google AI Studio, the Gemini API, and Android Studio, and Google upgraded its Gemini Spark personal assistant, available in 160+ countries, to run on 3.7 Flash. Google's own positioning is blunt: its "most intelligent workhorse model yet," built for complex coding, agentic workflows, and reliable multi-step execution. It arrives only about three weeks after Gemini 3.6 Flash — Google has compressed the Flash line to a near-monthly cadence, and this release is another confirmation that the Flash tier, not the flagship tier, is where Google is fighting for developers. This article breaks it down across specs, Google-reported benchmarks, pricing, and peer comparison, then answers the practical question: which mainland China relay providers already carry it, and how do you wire it up?

A note on fact discipline and the time stamp. This article is verified as of August 19, 2026. Google-side specs and pricing come from the official Gemini API documentation (model catalog and pricing page); Google-reported benchmarks are relayed from third-party coverage (such as Gadgets 360) of Google's own announcements; third-party independent evaluation (Artificial Analysis' Intelligence Index, AAII) has no credible score for Gemini 3.7 Flash on the public leaderboard as of publication — not because it has no score, but because none has been independently published yet. We say that explicitly rather than inventing a number to make a table look complete. Anywhere figures cannot be cross-checked, we say so. In particular, because 3.7 Flash is only days old, treat any "third-party benchmark says it beats X" claim with suspicion until you verify the source and the benchmark version.

1. Release & positioning: a three-week iteration, and the Flash tier as the main battleground

To understand Gemini 3.7 Flash, you first need to know where it sits in Google's lineup, because that sets the coordinates for every comparison below. Google's Gemini product line currently splits into roughly two tiers: Flash (the volume workhorse) and Pro (the flagship). On July 26, 2026, Google CEO Sundar Pichai told 9to5Google that the company is training a "much larger" next-generation base model, Gemini 4 (expected around November–December), while keeping the Flash line on a near-monthly release cadence focused on agentic coding. Gemini 3.7 Flash is that cadence in action: 3.6 Flash landed around late July, 3.7 Flash on August 14 — roughly three weeks apart.

The official positioning is worth pausing on. The Gemini API model catalog describes 3.7 Flash as "our latest and most capable Flash model, built for complex coding, agentic workflows, and reliable multi-step execution," and the launch coverage uses an even more emotive phrase — "the most intelligent workhorse model yet." In other words, Google is not trying to pass this off as a flagship substitute; it is openly treating "coding + agents + high-throughput volume" — the most crowded, highest-volume lane in the market — as the Flash tier's home turf. This is the same playbook as open-weight coding workhorses like DeepSeek V4 Flash and GLM-5.3: the flagship builds the brand, the Flash tier wins the developers and the volume.

One more piece of context matters: Google's flagship tier is in an awkward window right now. Gemini 3.5 Pro has slipped repeatedly since June, and as of publication (August 19) the official line is still "being tested with partners, launching as soon as it's ready," with no pricing and no date; the highest tier you can actually buy remains Gemini 3.1 Pro (Preview). With the flagship on hold, the Flash tier gets pushed further forward, and 3.7 Flash is effectively "the newest Google model you can actually buy." That matters for relay selection: in August 2026, if you want the latest Google model through a relay, the Flash tier is essentially the only game in town.

2. Specs: ~1M input context, 64K output, native multimodal input, thinking levels and the tool suite

The hard specs first (all from Google's official Gemini API model docs, verified as of publication):

SpecGemini 3.7 Flash
Model IDgemini-3.7-flash (Google also uses aliases like gemini-3.7-flash@stable)
Input context1,048,576 tokens (~1M)
Max output65,536 tokens
Input modalityText, image, video, audio, PDF (native multimodal input)
Output modalityText only (no audio/image generation, no Live API)
Thinkinglow / medium / high; minimal is not supported (errors)
Tools / capabilitiesFunction calling, code execution, caching, file search, structured outputs, URL context, Google Search/Maps grounding, Computer Use (preview)
Inference tiersBatch / Flex / Priority, with different price and speed

On specs, 3.7 Flash continues the Flash-tier formula: flagship-grade context, multimodal input, text-only output. The ~1M input context lets it swallow a whole repo or a long document in one pass; the 65,536 (64K) max output is enough for generating a full test suite or refactoring multiple files, but it is a real boundary compared with ultra-long-output models like DeepSeek V4 Pro (~384K). Multimodal input (text/image/video/audio/PDF) has long been a Flash-tier strength — in "read this screenshot / UI mock / scanned PDF" scenarios it beats many same-priced open-weight models — but because output is text-only, it cannot replace image/video generation models.

"Thinking" is one of the keys to understanding 3.7 Flash's cost: it supports low/medium/high reasoning intensity, and the official pricing includes thinking tokens in the output price ($3.75/M output already covers thinking, see Section 4) — unlike models that bill thinking separately or quietly raise the price when thinking is forced on. And "Computer Use (preview)" signals that Google is pushing "operate-a-computer" agents down into the Flash tier — a signal for automation/browser-agent developers, but since it's a preview, stability and whether your relay passes it through are things to verify yourself. Also note: the official docs say the minimal thinking level is not supported (it errors), so check that your calling pattern doesn't hit it.

3. Benchmarks: how to read DeepSWE going from 49.0% to 65.3%, and why there's no third-party AAII yet

Pay attention here, because this is the easiest place to get misled: the benchmarks below are all Google-reported (relayed from launch coverage), and as of publication there has been no large-scale third-party re-run. We present them as "vendor claims," with the source flagged, not as independent fact.

BenchmarkGemini 3.6 FlashGemini 3.7 FlashChange
DeepSWE v1.1 (software engineering)49.0%65.3%+16.3pp
FrontierCode 1.1 (coding)34.4%43.6%+9.2pp
AutomationBench (automation workflows)17.0%30.4%+13.4pp

There are three ways to read these three numbers. First, the absolute jumps are large: DeepSWE v1.1 used to be near the "passing line" for most models a year ago, and 3.7 Flash moves from 49.0% to 65.3%, while AutomationBench nearly doubles (17.0%→30.4%). This is not incremental — it's a real post-training step up. Second, it now sits alongside same-tier closed models on coding benchmarks: at 43.6% on FrontierCode 1.1 it narrowly beats Claude Sonnet 5's 42.7% and GPT-5.6 Terra's 41.3% in Google's own reporting — note "narrowly beats same-tier models," not "beats the flagships." Third, these are vendor-reported numbers, and each vendor's methodology differs, so treat cross-vendor comparisons cautiously.

So what does the third-party picture look like? As of publication, there is no credible AAII score for Gemini 3.7 Flash on Artificial Analysis' leaderboard — it's only days old and independent evaluation takes time. For reference, the current AAII front tier (third-party, already verified by this site) is: Claude Opus 5 (max) at 63, Claude Fable 5 at 62, GPT-5.6 Sol (max) at 61, Kimi K3 (max) at 60 (top open-weights); DeepSeek V4 Pro 0813 sits at 53. In other words, even if 3.7 Flash's vendor-reported numbers look strong, on a unified third-party scale it will most likely land in the workhorse tier, with a real gap to flagships like Opus 5 / Fable 5 / GPT-5.6 Sol — that's what the Flash tier is, not a defect of 3.7 Flash. Once Artificial Analysis or LMSYS publish updated numbers, this article's judgement can be checked against third-party data; until then, treat any "3.7 Flash beats a flagship" claim as unverified vendor marketing.

How to read this round of benchmarks

The vendor-reported DeepSWE 65.3 and FrontierCode 43.6 tell you "3.7 Flash is a large step up from 3.6 Flash and catches same-tier closed models" — a credible direction. But they don't replace unified third-party measurement. Today's AAII top five are still flagships or near-flagships like Opus 5 / Fable 5 / GPT-5.6 Sol / Kimi K3. Keep "vendor-reported" and "third-party" as two separate datasets and don't mix them.

4. Pricing: the $0.75/$3.75 six-month promo, the January doubling, and the "whole-Flash-line promo" truth

Pricing is the part worth the most attention, because it hides a detail most coverage got wrong. Here's the official rate card (Gemini API pricing page, Standard tier, per million tokens):

ItemPromo (through 2026-12-31)Standard (from 2027-01-01)
Input$0.75$1.50
Output (incl. thinking)$3.75$7.50
Context caching (hit)$0.075$0.15
Caching storage$0.50 / 1M tokens/hour$1.00 / 1M tokens/hour

Two details worth spelling out. First, the "half-price promo" is not exclusive to 3.7 Flash — it's a whole-Flash-line promo shared with 3.6 Flash: on the official pricing page gemini-3.6-flash has the same $0.75/$3.75 promo and the same $1.50/$7.50 reversion in January. So "input at $0.75, half the previous generation's price" is Google opening a discount window across the Flash volume line to pull developers in before year-end; if you read this after December 31, 2026, both 3.7 and 3.6 will have reverted. Second, 3.7 Flash's promo output price of $3.75 already includes thinking tokens — an important difference from models that bill thinking separately or quietly raise the effective price — so you don't need to add a separate thinking surcharge when estimating cost.

There's also a pricing quirk worth noting: Gemini 3.5 Flash (the previous generation) is listed at $1.50 input / $9.00 output — its output is more expensive than 3.6/3.7's standard $7.50. In other words, Google has been actively pushing the Flash tier's output unit price down across 3.5 → 3.6 → 3.7, consistent with the volume-workhorse positioning and with the pricing patterns this site documented in its August 2026 pricing roundup (Batch API at 50% off, cache reads at ~10% of miss price). On the flagship side: Gemini 3.1 Pro (Preview) remains $2/$12 under 200K context and $4/$18 above; 3.5 Pro is unreleased with no pricing; Gemini 4 is still in training.

A concrete cost example makes "$0.75/M" tangible. Say you run a whole-repo refactor over a 200K-token codebase with ~10K tokens of output (during the promo): input ≈ $0.75 × 0.2 ≈ $0.15, output ≈ $3.75 × 0.01 ≈ $0.038, so under $0.19 per task; if the codebase prefix is reused heavily, the $0.075/M cache-hit input price pushes it lower. For comparison, the same task on Claude Sonnet 5 (promo $2/$10) is roughly $0.40 + $0.10 = $0.50, and on GPT-5.6 Terra ($2/$12) roughly $0.40 + $0.12 = $0.52 — so during the promo, 3.7 Flash lands around a third of those two same-tier rivals. That, not the benchmarks, is the real point of "volume pricing."

5. Peer comparison: Flash tier vs flagship tier, and who's actually cheaper per token

Placing 3.7 Flash in the August 2026 model price table (all first-party official prices, promo prices marked with their expiry; DeepSeek/GLM converted to USD):

ModelContextInput $/MCache hit $/MOutput $/M
Gemini 3.7 Flash (promo to 12/31)~1M$0.75$0.075$3.75 (incl. thinking)
Gemini 3.7 Flash (standard, from Jan)~1M$1.50$0.15$7.50
Gemini 3.6 Flash (promo to 12/31)$0.75$0.075$3.75
Gemini 3.5 Flash (official, no promo)$1.50$0.15$9.00
Gemini 3.1 Pro (Preview, ≤200K)200K tier$2.00$0.20$12.00
Claude Sonnet 5 (promo to 8/31)1M$2.00 (std $3.00)$10.00 (std $15.00)
GPT-5.6 Terra1.05M$2.00read $0.20$12.00
DeepSeek V4 Flash (0731)1M$0.14≈$0.0028$0.28
DeepSeek V4 Pro (0813)1M$0.435≈$0.0036$0.87
Kimi K31M$3.00$0.30$15.00
GLM-5.3 (API not live; 5.2 pricing as reference)1M~$1.40~$0.26~$4.40
Qwen3.8-Max1M$2.00$6.00

Three takeaways. First, during the promo, 3.7 Flash is the cheapest tier among closed same-tier models: input at $0.75 is roughly a third of Claude Sonnet 5 (promo $2) and GPT-5.6 Terra ($2), and output at $3.75 is roughly a third of theirs too; versus Google's own 3.1 Pro ($2/$12), the promo saves about two-thirds. Second, it still loses to the open-weight volume kings: DeepSeek V4 Flash ($0.14/$0.28) and V4 Pro ($0.435/$0.87) remain the absolute lowest-priced tier — 3.7 Flash's promo input is over 5x V4 Flash's and output over 13x. If "extremely cheap + direct from mainland China" is your first requirement, DeepSeek remains more extreme; 3.7 Flash's differentiation is Google's ecosystem, native multimodal input, and near-same-tier coding benchmarks. Third, don't ignore the January reversion: once the promo ends, 3.7 Flash returns to $1.50/$7.50, where its per-token price advantage over Sonnet 5 / Terra basically evaporates (input parity, output still a bit lower), leaving ecosystem and multimodal differences to carry it — so "use the promo price before year-end" isn't marketing; it's structural.

On positioning, draw an honest line: 3.7 Flash is a workhorse-tier model, not a flagship. Its rivals are Claude Sonnet 5, GPT-5.6 Terra, DeepSeek V4 Pro, GLM-5.3, and Kimi K3. Comparing it to Claude Opus 5 ($5/$25), GPT-5.6 Sol ($5/$30), or Claude Fable 5 ($10/$50) is a flagship-tier conversation the Flash tier doesn't participate in. If you want "cheap and Google," 3.7 Flash is the sweet spot among closed models during the promo; if you want the most intelligence on earth, go look at the flagships — don't wave a Flash model's vendor-reported score at a flagship.

6. Reaching it via relay providers: who already carries gemini-3.7-flash, and the standard access pattern

For mainland-China readers, the first hurdle to Gemini 3.7 Flash isn't price — it's the network: Google's Gemini API (ai.google.dev / generativelanguage.googleapis.com) is not directly reachable from the mainland, the official channel needs a proxy, and it doesn't accept CNY top-ups. That is exactly the core value of a mainland China AI API relay: a mainland-reachable base_url plus a CNY-funded key that "imports" overseas models. So "how do I reach 3.7 Flash via a relay" is really two questions: which relay already carries the model ID, and what does the integration code look like.

From the providers in this site's database that were recently verified (lastVerified 2026-08-16), several already list Gemini 3.7 in their coverage, and ProAI API (proaiapi.tech) explicitly notes gemini-3.7-flash is live (its pricing note reads "fast roll-out of new models (grok-4.6/gemini-3.7-flash/qwen3.8-max live)"); MKEAI's tagline and coverage mention Gemini 3.7; NoDAPI (545+ models) and JENIYA (454+ models) also list Gemini 3.7/3.6/3.5. A necessary caveat: these are the providers' own coverage claims, and "covers Gemini 3.7" does not guarantee every feature (especially Computer Use preview and the thinking levels) is passed through stably. Check the actual model list and price in each provider's console, and test with a small top-up before committing.

The integration pattern is identical to every OpenAI-compatible relay this site has covered — just point at gemini-3.7-flash. With the OpenAI SDK:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_RELAY_API_KEY",
    base_url="https://your-relay-domain/v1"  # use the actual base_url from your relay's console
)

response = client.chat.completions.create(
    model="gemini-3.7-flash",
    messages=[
        {"role": "system", "content": "You are a senior full-stack engineer"},
        {"role": "user", "content": "Do a whole-repo refactor of this code and add tests"},
    ],
)
print(response.choices[0].message.content)

Through LiteLLM or a gateway, it's the same idea:

model_list:
  - model_name: gemini-3.7-flash
    litellm_params:
      model: gemini/gemini-3.7-flash
      api_key: os.environ/GEMINI_API_KEY

Important: gemini-3.7-flash is the official/common-gateway convention, not a guarantee every relay has it live. Some relays expose google/gemini-3.7-flash or gemini-3.7-flash@stable; since 3.7 Flash is days old, whether each relay has followed up and how they price it (the official promo $0.75/$3.75 usually gets a FX margin on top at relays) is something you must verify console-by-console. Don't point production traffic at an ID just because this article wrote it down. Also note: Google's Batch API is 50% off across the line — if your relay supports Batch/Flex tiers, batch jobs get even cheaper.

If you're using AI coding tools like Claude Code, Cursor, or Gemini CLI, the pattern is "swap base_url + key": Claude Code points ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN at a relay that speaks the Anthropic format; Cursor sets the model provider's base_url to the relay and fills gemini-3.7-flash as the model name; Gemini CLI supports GEMINI_API_KEY with a custom endpoint. Whether each tool is passed through depends on the relay's protocol support — not every relay forwards all three tools the same way, so validate on one tool with small traffic before production.

7. The verdict: when to use it, and when not to

  • Mainland developer + want the newest Google model + price-sensitive → during the promo, go straight to gemini-3.7-flash. It's the newest Google model you can actually buy through a relay (flagship 3.5 Pro still unreleased), and $0.75/$3.75 is the cheapest tier among closed same-tier rivals. Start with a small top-up and test.
  • Coding / agents / multi-step tool use → this is its home turf: ~1M context, function calling, code execution, search grounding, Computer Use (preview). But Computer Use is a preview — verify your relay passes it through before production.
  • Multimodal input (screenshots / UI / PDFs) → the Flash tier's edge over same-price open-weight models; just remember output is text-only.
  • Extreme cheap + direct from mainland + ecosystem-agnostic → DeepSeek V4 Flash/Pro remains the absolute lowest-priced tier; 3.7 Flash doesn't beat it on pure price. For pure-text, tight-budget workloads, DeepSeek is more extreme.
  • Want the most intelligence / flagship benchmarks → look at Claude Opus 5 / GPT-5.6 Sol / Fable 5. Don't compare a Flash-tier model to flagships.
  • Long-document heavy + batch workloads → 3.7 Flash's $0.075/M cache-hit price (promo) is low for a closed model and Batch is 50% off; but if you have extreme hit rates and want maximum savings, DeepSeek's ~$0.0028/M cache price is still an order of magnitude cheaper.

8. Conclusion: what 3.7 Flash actually signals

Gemini 3.7 Flash is a "quiet launch, loud volume play." There was no event, no "beats the flagship" press release — just a solidly specified model with big vendor-reported benchmark gains and an aggressive $0.75/M price, shipped straight into production as the engine of Gemini Spark across 160+ countries. Three signals are worth remembering.

First, Google has made the Flash tier its main battlefield. With flagship 3.5 Pro repeatedly delayed and Gemini 4 still in training, the newest Google model you can buy today is Flash-tier — 3.6 and 3.7 shipped within a month of each other, and the near-monthly cadence is aimed squarely at the "coding + agents + volume" lane. Second, "$0.75/$3.75" is a six-month promo across the whole Flash line, not exclusive to 3.7. 3.6 Flash is promo-priced identically through December 31, 2026, and both revert to $1.50/$7.50 in January. If you want the promo price, use it this year; if you're estimating long-term cost, price it at standard rates. Third, it resets the price reference for the closed "workhorse" tier. During the promo its per-token cost is about a third of Claude Sonnet 5 / GPT-5.6 Terra, pushing the "closed can be cheap too" narrative a notch lower — even though it still loses to open-weight volume kings like DeepSeek, for mainland developers who want both the Google ecosystem and a controlled budget, this is the smoothest Google entry point in August 2026.

One last reminder for relay readers: 3.7 Flash is days old, third-party benchmarks (AAII) are not out yet, and whether each relay passes it through stably is unverified. The specs and official rate card above are verified as of August 19; the vendor-reported benchmarks are "vendor claims, relayed with skepticism"; third-party numbers should be checked after publication. Before integrating, open your chosen relay's console and confirm gemini-3.7-flash is actually in the model list, then test with a small top-up — that's more reliable than any article.

Putting it together

  • Want the newest Google model + price-sensitive → during the promo, integrate gemini-3.7-flash; several relays already carry it (ProAI API explicitly live; MKEAI/NoDAPI/JENIYA list Gemini 3.7).
  • Coding / agents / multimodal input → its home turf, ~1M context + full tool suite; but text-only output and Computer Use is still preview.
  • Pure-text maximum savings → DeepSeek V4 remains the absolute lowest price; 3.7 Flash promo is ~5x V4 Flash on input.
  • One timing reminder: the $0.75/$3.75 promo ends 2026-12-31, both 3.6/3.7 revert to $1.50/$7.50 in January; third-party AAII and relay pass-through quality should be re-checked after publication.