Here's the conclusion up front, so you don't have to scroll: as of September 18, 2026, neither Gemini 4 nor Gemini 4 Pro has been released, and neither has official confirmed specs, pricing, or a launch date. In Google's official Gemini API model catalog (page last updated 2026-09-17), the highest number is Gemini 3.8 Flash, which hit GA on September 2. The flagship Pro tier is still Gemini 3.1 Pro, released February 19, 2026 and still carrying a Preview label. As for Gemini 3.5 Pro, which slipped from its June target, the official page still says "coming soon." In other words, if you want to "use Google's newest and strongest Gemini" today, your ceiling is the 3.8 Flash + 3.1 Pro combination — not anything called "4." This article breaks that lineup down along four real axes: real specs, real prices, real timing, and real access routes. It also answers the question our readers actually care about: which mainland China relay providers can reach it, and under which model ID.

1. Where Gemini 4 Pro actually stands: what Google said, and why the "October launch" claim doesn't hold up

This has to come first, because the signal-to-noise ratio around Gemini 4 is terrible — and this site's rule is simple: we publish what we can verify, not what sounds exciting.

We checked all three of Google's official outlets. First, the Gemini API model catalog (ai.google.dev/gemini-api/docs/models, page marked last updated 2026-09-17): the complete model list contains no Gemini 4, no Gemini 4 Pro, and nothing at all numbered 4. The banner at the top reads "Gemini 3.8 Flash is now available." Second, the official Gemini API changelog: across every 2026 entry, the highest-numbered model is gemini-3.8-flash, GA on September 2. There is no Gemini 4 record of any kind. Third, Google DeepMind's Gemini model page: the cross-model comparison table again has no Gemini 4; the only forward-looking text is "3.5 Pro coming soon" — note, 3.5 Pro, not a fourth generation. On top of that, blog.google's product pages and its on-site search return zero hits for "Gemini 4."

So is the idea that Google is working on Gemini 4 true? Yes — but only as a single official statement. On July 26, 2026, Google CEO Sundar Pichai confirmed in an interview reported by 9to5Google that the company is training Gemini 4 with "much larger base models," aiming for it to be at the frontier "of where the frontier will be" at the moment it launches. He added that Google's "first priority" for TPU allocation is "making sure we are allocating what we need to compete at the frontier in terms of AGI development." That statement contains no date, no specs, and no pricing. The widely repeated "expected in November or December" is journalists extrapolating from Google's past cadence — not a Google commitment.

About the "Gemini 4 Pro launches in October" claim

The one article behind that specific month — in both English and Chinese circulation — is a September 9, 2026 piece on Geeky Gadgets, and it cites a YouTube video. The piece contains two hard errors: it spells Anthropic as "Enthropic," and it describes Grok as a SpaceX product. An outlet that gets a competitor's name and ownership wrong is not a reliable source for a launch timetable — and the claim has zero corroboration in Google's official channels. We therefore don't use it. To judge whether you should wait for Gemini 4, go by the official model catalog and changelog, not secondhand coverage.

2. Google's real Gemini lineup in one table

Rather than wishing on a model that doesn't exist, it helps to see the matrix Google is actually running. Everything below comes from Google's official model catalog and pricing page (verified 2026-09-17/18):

ModelModel IDStatusPositioning
Gemini 3.8 Flashgemini-3.8-flashStable (GA 2026-09-02)Newest, most capable Flash — long-horizon software engineering and agents
Gemini 3.8 Flash Cyber(not published)Limited (Fairwind Program)Cybersecurity: vulnerability discovery + automated patching
Gemini 3.8 Livegemini-3.8-liveStable (GA 2026-09-15)Default for low-latency voice agents
Gemini 3.8 Live Extended Thinkinggemini-3.8-live-extended-thinkingStable (GA 2026-09-15)Live audio plus background reasoning
Gemini 3.7 Flashgemini-3.7-flashStable (GA 2026-08-13)Previous-gen coding/agent workhorse
Gemini 3.6 Flashgemini-3.6-flashStablePrior Flash generation
Gemini 3.5 Flashgemini-3.5-flashStable (legacy)Routine high-throughput work
Gemini 3.5 Flash-Litegemini-3.5-flash-liteStableFastest, cheapest 3.5 model
Gemini 3.1 Flash-Litegemini-3.1-flash-liteStableFrontier-class performance at a fraction of the cost
Gemini 3.1 Progemini-3.1-pro-previewPreviewHighest purchasable Pro tier — complex reasoning, vibe coding
Gemini Omni Flashgemini-omni-1.1-flashStable (GA 2026-08-27)Video generation, editing, extension
Nano Banana Progemini-3-pro-imageStableTop-tier image generation and editing
Nano Banana 2gemini-3.1-flash-imageStableHigh-volume image generation
Veo 3.1 / Lyria 3.5veo-3.1-* / lyria-3.5Preview / StableVideo generation / music generation
Gemini 3.5 Pro—UnreleasedOfficial page says "coming soon"; still in partner testing
Gemini 4 / Gemini 4 Pro—Unreleased, no published specsGoogle has only confirmed it is "training" it

The most notable thing in that table is an unusually long seven-month gap in Google's Pro tier. Gemini 3.1 Pro shipped on February 19, 2026 and had not been superseded by any newer Pro model by mid-September — and it still carries a Preview label. Gemini 3.5 Pro, widely expected, has slipped repeatedly since June. The August 5 restructuring at Google DeepMind (Demis Hassabis moving to DeepMind chairman and Alphabet chief scientist, with former CTO Koray Kavukcuoglu stepping up as SVP reporting directly to Pichai and overseeing Gemini development) is another signal of the pressure Google has been under at the frontier. What has kept the product line alive is the Flash tier's roughly monthly cadence.

3. Gemini 3.8 Flash: GA on September 2, 2026 — official benchmarks and specs

This is Google's newest — and most worthwhile — Gemini to integrate today. The official changelog is explicit: on September 2, 2026, gemini-3.8-flash reached GA, described as "our most intelligent Flash model," aimed at long-horizon software engineering, agents, and enterprise workflows.

Specs (official model catalog): 1,048,576 input tokens (about 1M) and a maximum of 65,536 output tokens (64K); input modalities cover text, images, video, audio, and PDF, while output is text only; tool support includes function calling, search as a tool, and computer use. Availability spans the Gemini app, Gemini Enterprise Agent Platform, Google AI Studio, the Gemini API, Google AI Mode, and Google Antigravity. Its status is general availability — not a preview.

Official benchmarks (Google-reported, from the DeepMind model page):

BenchmarkGemini 3.8 FlashComparison
DeepSWE v1.1 (long-horizon software engineering)73.7%Google claims it beats most larger frontier models at a fraction of the cost per task
Terminal-bench 2.189.4%—
Terminal-bench 4.019.1%Harder next-gen version; scores are low across the industry
HLE-Verified (Humanity's Last Exam)54.9%GPT-5.6 Sol 54.5%, Claude Opus 5 54.4%, 3.7 Flash 53.6%
Vals Finance Agent v261.4%3.7 Flash 59.0%, Claude Opus 5 58.6%, GPT-5.6 Terra 54.4%
Harvey's Legal Agent (all-pass)10.0%Leads the next entry at 8.8%
OSWorld-2.0 (computer-use agent)59.0%—
GDPVal-AA v21545 Elo—
GDP.PDF / CharXiv Reasoning35.0% / 86.2%Document and chart understanding
LVBench (long video, agentic)87.8%—
LABBench2 / BioMysteryBench86.2% / 88.8% (human-solvable)Scientific and research scenarios

How should you read that set? First, coding and terminal ability is the hardest improvement in this generation: DeepSWE v1.1 is pushed to 73.7%, and Terminal-bench 2.1 at 89.4% is a strong position for the tier. Second, on two specialist agent benchmarks it outright beats larger rivals — 61.4% on Vals Finance Agent v2 is ahead of Claude Opus 5's 58.6%, and its 10.0% on Harvey's legal agent ranks first. Third — and this is where restraint matters — every number above is Google-reported. Vendor methodologies differ, so leave headroom in any cross-vendor comparison. Treating "vendor-reported" and "third-party leaderboards" as two separate data sets is this site's standing advice.

One more update that's easy to miss: on September 1, 2026, Google shipped agentic video understanding for 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, claiming up to 88% fewer tokens on long-form video content. For anyone building video-content analysis or long meeting summaries, that's a genuine cost optimization — and it was pushed down as far as the cheapest tier, 3.5 Flash-Lite.

4. Gemini 3.8 Flash Cyber: vulnerability discovery and auto-patching, but not sold publicly

This is one of Google's most earnest bets on vertical models in 2026, and one of the least clearly explained in Chinese-language coverage. DeepMind describes it as "our most capable cybersecurity model," with frontier-level vulnerability detection and automated patching. It's built on the Flash line, so it keeps that speed and low cost, which "enables quick iteration."

Functionally, it works through large codebases spanning 20 programming languages to surface hidden flaws on its own, then generates validated, high-quality fixes automatically. Google positions it as "designed specifically for defenders." The reported benchmarks:

BenchmarkGemini 3.8 Flash CyberComparison
CyberGym Pass@1 (vulnerability discovery)86.2%GPT-5.5-Cyber 85.6%, Mythos 5 83.8%, GPT-5.6 Sol 83.6%, 3.5 Flash Cyber 77.5%
Internal real-world discovery benchmark (20 languages)71.0%3.7 Flash 58.9%, 3.5 Flash Cyber 46.6%
CWE-Bench Pass@1 (external, run by Collinear)47.2%Just under Fable 5's 47.8%, but at roughly $3.60 per rollout — on the Pareto frontier
Gray Swan IPI (prompt-injection attack success rate, lower is better)6.0%—

The crucial constraint is access: Gemini 3.8 Flash Cyber is not publicly sold. It is available only "to a set of trusted defenders via the Fairwind Program," which Google describes as giving partners "a critical head start against AI-driven cyber threats." The official page gives no list price and no general-availability date. The practical takeaway for mainland China readers: it will not be purchasable through a relay provider anytime soon, and any relay claiming to offer "Gemini 3.8 Flash Cyber" deserves serious suspicion. If this capability class is what you need, Zhipu's GLM-5.3 (CyberGym 84.5%), which we reviewed separately, follows the same "trusted access" gating logic.

5. Gemini 3.8 Live and Live Extended Thinking: the voice-agent pair that went GA September 15

On September 15, 2026, Google brought two audio-to-audio Live API models to GA — less than two weeks after 3.8 Flash:

  • gemini-3.8-live — positioned as "the default option for most low-latency voice agent experiences," with no reasoning delay, optimized for immediate response;
  • gemini-3.8-live-extended-thinking — supports background reasoning alongside live audio, for complex scenarios where the agent needs to think before answering.

Both have a free tier. Paid input runs $0.75/M for text, $3.00/M for audio (or $0.005 per minute), and $1.00/M for image and video (or $0.002 per minute); output runs $4.50/M for text and $12.00/M for audio (or $0.018 per minute). That per-minute structure signals where Google wants this deployed: phone support, real-time translation, and always-on voice assistants — businesses billed by call duration. For reference, Google now recommends migrating off the 3.1 Flash Live Preview.

If your product involves real-time voice, this is Google's only current recommended integration point. Note, however, that the Live API uses a WebSocket long connection, which differs substantially from a conventional HTTP relay's compatibility surface — confirm before production whether your chosen relay supports streaming audio passthrough, since many only proxy text chat completions.

6. Gemini 3.1 Pro: the only purchasable Pro tier, and why it's still stuck on Preview

If what you actually need is "Google's flagship intelligence," the answer is not Gemini 4 Pro — it is Gemini 3.1 Pro, live for over seven months but still in Preview (model ID gemini-3.1-pro-preview, released February 19, 2026).

Specs: 1M input tokens and 64K maximum output; input supports text, images, video, audio and PDF with text-only output; tools include function calling, structured output, search as a tool, and code execution. Google positions it as "our most intelligent model yet," best for complex tasks and deep work, highly capable at vibe coding and agentic coding, with better tool use and simultaneous multi-step tasks. Availability matches 3.8 Flash (Gemini app, AI Studio, Gemini API, Gemini Enterprise Agent Platform, AI Mode, Antigravity).

Official benchmarks: 44.4% on Humanity's Last Exam (no tools), 77.1% on ARC-AGI-2 (ARC Prize Verified), 94.3% on GPQA Diamond, 68.5% on Terminal-Bench 2.0, 80.6% on SWE-Bench Verified (just behind Opus 4.6's 80.8%), 2887 Elo on LiveCodeBench Pro, 92.6% on MMMLU, and 84.9% on MRCR v2 at 128k.

There's an easily missed comparison buried in those numbers: 3.1 Pro scores 80.6% on SWE-Bench Verified, while 3.8 Flash scores 73.7% on the newer DeepSWE v1.1 — the two benchmarks aren't directly comparable, but the direction is clear. Google has moved fast at closing the gap between its workhorse and flagship tiers, and slowly at replacing the flagship itself. That is precisely the cost of Gemini 3.5 Pro's repeated delay.

On 3.5 Pro: the official line remains that Google is testing it with partners and will launch "as soon as it's ready." DeepMind's model page still reads "3.5 Pro coming soon," with no pricing and no date. It was originally targeted at June and is now months behind. As for Gemini 4 — go back to section 1. All Google has said is that it is training it.

7. Gemini Omni 1.1 Flash: GA August 27, and video generation now does extension and keyframe interpolation

If the models above are the text line, Omni 1.1 Flash is Google's newest move in generative video. On August 27, 2026, gemini-omni-1.1-flash reached GA. Compared with gemini-omni-flash-preview, which entered public preview on June 30, it adds three things:

  • Video extension: via the extend task, each call adds 10 seconds, stacking up to a continuous 40-second video — the first time the API has supported this;
  • First/last-frame interpolation: supply a starting frame and an ending frame and the model generates the intervening video;
  • A resolution parameter: supporting 360p, 720p (default), 1080p, and 4K. Google also offers a cheaper 360p draft mode that it says delivers up to 60% more throughput at roughly a third of the 720p cost. Note that 1080p and 4K are upsampled output rather than native resolution.

On pricing, Omni Flash input is $1.50/M (text, image, video, audio), with output at $9.00/M for text and $17.50/M for video. Google's pricing page also lists per-second rates: $0.03/sec at 360p, $0.10/sec at 720p, $0.15/sec at 1080p, and $0.30/sec at 4K. There is no free tier, and context caching is not supported. The older gemini-omni-flash-preview endpoint is scheduled for deprecation on September 30, 2026, so anyone still on the preview needs to migrate promptly.

8. Official pricing, itemized: promo vs standard, and the Batch / Priority tiers

All prices below come from Google's official Gemini API pricing page (verified 2026-09-18), in USD per million tokens. The single most important thing to remember: the Flash line's promo pricing ends December 31, 2026, and doubles on January 1, 2027.

ModelInput (Standard / Batch·Flex / Priority)Output (Standard / Batch·Flex / Priority)Free tier
Gemini 3.8 Flash$0.75→$1.50 / $0.375→$0.75 / $1.35→$2.70$3.75→$7.50 / $1.875→$3.75 / $6.75→$13.50Yes
Gemini 3.7 FlashSame as aboveSame as aboveYes
Gemini 3.6 FlashSame as aboveSame as aboveYes
Gemini 3.5 Flash$1.50 / $0.75 / $2.70$9.00 / $4.50 / $16.20Yes
Gemini 3.5 Flash-Lite$0.30 / $0.15 / $0.54$2.50 / $1.25 / $4.50Yes
Gemini 3.1 Flash-Lite$0.25 / $0.125 / $0.45$1.50 / $0.75 / $2.70Yes
Gemini 3.1 Pro (Preview)$2.00 (≤200K) / $4.00 (>200K)$12.00 (≤200K) / $18.00 (>200K)No
Gemini Omni Flash$1.50$9.00 text / $17.50 videoNo
Nano Banana Pro (image)$2.00$12.00 text / $120.00 imagesNo
Nano Banana 2 (image)$0.50$3.00 text / $60.00 imagesNo
Gemini 3 Flash (Preview)$0.50 (text/image/video), $1.00 audio$3.00Yes
Gemini 2.5 Pro$1.25 (≤200K) / $2.50 (>200K)$10.00 / $15.00Yes

Three pricing details worth calling out. First, "$0.75/$3.75" is not exclusive to 3.8 Flash — it's a promo across the whole Flash volume line. Three generations (3.6, 3.7, 3.8) share the identical promo and all revert to $1.50/$7.50 after 2026-12-31. Google is using a half-price window to pull developer traffic in, not giving one generation special treatment. Second, the Flash output promo price already includes thinking tokens, so you don't need to add a separate reasoning surcharge when estimating — a meaningful difference from models that bill thinking separately. Third, the Pro tier and the video/image models have no free tier, so trying 3.1 Pro requires funding an account, and the free Google Search grounding allowance of 5,000 requests per month is shared across all Gemini 3.x models, then $14 per 1,000 requests.

A quick worked example to make "$0.75/M" concrete. Say you're refactoring a 200,000-token codebase with roughly 10,000 output tokens, on 3.8 Flash during the promo: input costs about $0.75 × 0.2 ≈ $0.15 and output about $3.75 × 0.01 ≈ $0.038 — under $0.19 per task. Route it through Batch and both sides halve again, to roughly $0.094. The same task on 3.1 Pro (≤200K tier) runs $2.00 × 0.2 = $0.40 input plus $12.00 × 0.01 = $0.12 output, about $0.52 — 3.8 Flash during the promo is roughly a third of that. Once the promo ends in January 2027 and Flash returns to $1.50/$7.50, that gap narrows to about 60%. If you want the promo price, use it this year; if you're modeling long-run costs, model the standard price.

9. The complete 2026 release timeline: what cadence Google is actually running

Line up the dates from the official changelog and the year's rhythm becomes obvious:

DateEvent
2026-02-19gemini-3.1-pro-preview released (still the only purchasable Pro tier, still Preview)
2026-03-09Gemini 3 Pro Preview shut down; gemini-3-pro-preview now points to 3.1 Pro
2026-06-30gemini-omni-flash-preview enters public preview (3–10 second, 720p video generation)
2026-07-26Pichai confirms Google is training Gemini 4 ("much larger base models," no date)
2026-08-05DeepMind leadership change: Hassabis to chairman, Kavukcuoglu overseeing Gemini; chief scientist Jeff Dean departs
2026-08-13gemini-3.7-flash reaches GA
2026-08-27gemini-omni-1.1-flash reaches GA (video extension, keyframe interpolation, resolution parameter)
2026-09-01Agentic video understanding ships (3.7 Flash, 3.6 Flash, 3.5 Flash-Lite; up to 88% fewer tokens)
2026-09-02gemini-3.8-flash reaches GA
2026-09-03Lyria 3.5 reaches GA
2026-09-15gemini-3.8-live and gemini-3.8-live-extended-thinking reach GA
2026-09-17antigravity-preview-09-2026 replaces the 05-2026 version; official model catalog updated

That table is itself the strongest answer to "is Gemini 4 coming?" Between August and September 2026, Google shipped five things — 3.7 Flash, Omni 1.1 Flash, agentic video understanding, 3.8 Flash, and the two 3.8 Live models — and every one of them belongs to the 3.x generation. Not a single one is a fourth generation. The Flash line's iteration cycle has compressed to roughly three weeks, while the Pro line has gone seven months without an update and 3.5 Pro still hasn't landed. The reasonable inference: Google has bet its flagship transition on Gemini 3.5 Pro and Gemini 4, and in the meantime is holding market presence through rapid Flash iteration. Whether Gemini 4 appears in November or December, Google hasn't said — and we won't guess.

10. How to reach it: the official API barrier and the mainland China relay situation

For mainland China readers, the first obstacle to Gemini has never been price — it's network access. Google's Gemini API (ai.google.dev / generativelanguage.googleapis.com) cannot be reached directly from the mainland; the official channel requires a VPN and generally does not accept RMB top-ups. That is precisely the value of mainland AI API relay providers: a domestically reachable base_url plus an RMB-funded key, bringing overseas models like Google's back within reach.

First, the most honest statement available: as of September 18, 2026, none of the 153 relay providers cataloged on this site explicitly lists gemini-3.8-flash in its model lineup. The reason is straightforward — 3.8 Flash only reached GA on September 2, while our provider records were last verified on 2026-08-16, before it existed. That doesn't mean it will never be available; it means don't expect any article to hand you a verified 3.8 availability list right now. Check the console yourself before funding an account.

Here is the verifiable Gemini coverage as it stands:

  • Gemini 3.7 Flash (the newest Flash actually purchasable through relays today): ProAI API (proaiapi.tech) explicitly states in its pricing notes that "gemini-3.7-flash is live" — the only provider naming the specific model ID in its official description. MKEAI (tagline references Gemini 3.7; coverage spans Gemini 3.5–3.7; note that registrations for new users are currently closed), NodAPI (545+ models, covering Gemini 3.7/3.6/3.5), and Quanzil (454+ models, covering Gemini 3.7/3.6/3.5) all list the Gemini 3.7 series.
  • Gemini 3.1 Pro (the highest purchasable Pro tier): coverage is considerably broader — fourteen providers list gemini-3.1-pro in their supportedModels: ProAI API, Shenma Relay API, Yiye Zhiqiu API, No.1-API, TokenMix, UniAPI, AIFast.club, V-API, XycAi (Xingdao Intelligence), NoneLinear, Zhihui API, Nio API, XJAI (Hunter API), and Poe API.
  • Gemini 3.8 Flash / 3.8 Live / Omni 1.1 Flash / 3.8 Flash Cyber: no cataloged provider explicitly lists any of these yet. Because 3.8 Flash Cyber runs through the limited Fairwind Program, it is very unlikely to appear at a relay provider in the near term — treat any listing as suspect.

The integration pattern matches every OpenAI-compatible relay we've written about before: point at the base_url and swap in the Gemini model ID. With the OpenAI SDK:

from openai import OpenAI

client = OpenAI(
    api_key="your-relay-api-key",
    base_url="https://your-relay-domain/v1"  # use whatever your relay's console lists
)

# Workhorse tier: the newest Flash purchasable through relays today
response = client.chat.completions.create(
    model="gemini-3.7-flash",
    messages=[
        {"role": "system", "content": "You are a senior full-stack engineer."},
        {"role": "user", "content": "Refactor this repo end to end and complete the tests."},
    ],
)
print(response.choices[0].message.content)

# Flagship tier: the highest Pro model currently on sale
pro = client.chat.completions.create(
    model="gemini-3.1-pro-preview",
    messages=[{"role": "user", "content": "Review this research report deeply and argue against it."}],
)
print(pro.choices[0].message.content)

If you use a gateway layer like LiteLLM, the logic is the same — just point model at it:

model_list:
  - model_name: gemini-3.7-flash
    litellm_params:
      model: gemini/gemini-3.7-flash
      api_key: os.environ/GEMINI_API_KEY
  - model_name: gemini-3.1-pro
    litellm_params:
      model: gemini/gemini-3.1-pro-preview
      api_key: os.environ/GEMINI_API_KEY

Three cautions. First, gemini-3.7-flash is the conventional official spelling, but that doesn't mean every relay has it live; some may use variants like google/gemini-3.7-flash. Go by what your console actually lists. Second, relays typically layer exchange-rate spreads and markups on top of official pricing — Google's $0.75/$3.75 promo is first-party pricing, and the RMB figure you see at a relay won't necessarily scale proportionally. Third, if you use AI coding tools like Claude Code, Cursor, or Gemini CLI, the pattern is again "change the base_url, swap the key": Claude Code uses ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN (requiring Anthropic-compatible support on the relay), Cursor takes a base_url and model name in its provider settings, and Gemini CLI officially supports GEMINI_API_KEY with a custom endpoint. Not every relay proxies the same protocol to all three tools, so validate with low-volume traffic before production.

11. The verdict and conclusion

  • You're waiting for Gemini 4 Pro → Stop. It doesn't exist and has no release date. The only official statement is Pichai's July 26 remark about training "much larger base models." Every specific month you've seen circulating currently has no official backing.
  • You want Google's newest, strongest, actually purchasable model → Reach gemini-3.7-flash through a relay (ProAI API explicitly carries it; MKEAI, NodAPI and Quanzil cover the Gemini 3.7 series). 3.8 Flash is stronger (DeepSWE v1.1 73.7%), but no cataloged relay lists it yet — check the console yourself.
  • You need flagship-grade reasoning and deep analysis → Use gemini-3.1-pro-preview, covered by fourteen relays. At $2/$12 (≤200K) it costs roughly 3x the Flash promo and has no free tier. Remember it is still Preview, not a stable release.
  • You're doing coding, agents, or multi-step tool calling → This is the Flash line's home turf. Roughly 1M context swallows whole codebases, and function calling, search grounding and computer use are all available. Both 3.7 Flash (August 13) and 3.8 Flash (September 2) sit in this tier.
  • You're building real-time voice → gemini-3.8-live (low latency) or gemini-3.8-live-extended-thinking (with background reasoning), GA September 15. But confirm first whether your relay supports Live API streaming audio passthrough.
  • You're doing video generation → gemini-omni-1.1-flash, with extension and keyframe interpolation. Note the old preview endpoint is deprecated September 30, and per-second pricing (720p at $0.10/sec) adds up.
  • A timing reminder: the $0.75/$3.75 promo across the whole Flash line (3.6/3.7/3.8) ends 2026-12-31, reverting to $1.50/$7.50 on January 1, 2027. If you want the promo price, take it this year.

Look at Google in September 2026 and you find something slightly counterintuitive: nearly all of the company's most important product activity is happening in the Flash tier, not the flagship tier. 3.8 Flash on September 2, the two 3.8 Live models on September 15, Omni 1.1 Flash on August 27 — five models online within three weeks, all in the 3.x generation, none numbered 4. Meanwhile, in the Pro tier, 3.1 Pro hasn't been superseded since February, 3.5 Pro keeps slipping, and Gemini 4 amounts to one sentence from Pichai about training.

This isn't Google hiding cards; it's the real cycle of frontier model development. A base-model transition is far more expensive and slower than post-training iteration, and the Flash line's near-monthly cadence is exactly what sustains developer stickiness during a flagship gap while contesting the highest-volume arena: coding and agents. The practical implication for mainland China developers is clear: you don't need to wait for Gemini 4 Pro — you have two usable Gemini tiers right now, the workhorse 3.7 Flash (reachable through relays) and the flagship 3.1 Pro (reachable through fourteen providers). 3.8 Flash is stronger still, but you'll need to check your own console for availability. As for when Gemini 4 arrives, this article's stance is simple: we write it when Google says it, and until then we won't invent a single month.

Put together, this means

  • Gemini 4 Pro does not exist. Neither the official model catalog (updated 2026-09-17) nor the changelog contains a Gemini 4 entry. The only official statement is Pichai's July 26 remark about training "much larger base models" — no date, no specs, no price.
  • What's actually on sale: the workhorse tier gemini-3.8-flash (GA September 2; DeepSWE v1.1 73.7%, ~1M context, $0.75/$3.75 promo through 12/31) and the flagship tier gemini-3.1-pro-preview ($2/$12 and up; HLE 44.4%, SWE-Bench Verified 80.6%; still Preview).
  • Relay access: ProAI API explicitly carries gemini-3.7-flash; MKEAI, NodAPI and Quanzil cover the Gemini 3.7 series; fourteen cover gemini-3.1-pro; none lists gemini-3.8-flash — verify in the console before funding.
  • A timing reminder: the Flash line promo ends 2026-12-31, reverting to $1.50/$7.50 in January; the Omni Flash preview endpoint is deprecated September 30.