The short version: the way relays used to make money is being copied by the vendors themselves. A relay's core pitch has always been "get the same class of model for less." In September 2026, OpenAI turned that pitch into its own official pricing: the previous flagship Astra was $10/$50; GPT-6 Sol, shipped September 22, is $2/$10 and officially declared permanent with no expiry date; the light tier GPT-6 Luna is $0.10/$0.50. From Astra to Luna, output pricing fell by a factor of 100. That's good news for users. For a relay whose only differentiator is the spread, it's a question about the foundation of the business. In the model below, the same 20% markup leaves roughly ¥71 of gross spread per million output tokens in the Astra era, about ¥14 at Sol pricing, and about ¥0.7 at Luna pricing — while servers, bandwidth, payment rails, support and risk control don't get any cheaper at all. This isn't a eulogy. It's the arithmetic, laid out. And once it's laid out, you can see that relays still have very hard things to sell. Those things were simply never price.

1. Three numbers side by side: $10/$50 → $2/$10 → $0.10/$0.50

The entire argument of this article compresses into one row of numbers. Everything below comes from OpenAI's official pricing, in dollars per million tokens, ordered by release date:

ModelReleasedInput / OutputCacheOfficial framing
GPT-6 Astra2026-09-04$10 / $50—Previous-generation flagship pricing; ARC-AGI-3 99.9%, GPQA Diamond 96%
GPT-6 Sol2026-09-22$2 / $10—Halved from GPT-5.6 Sol's $4/$20; officially declared permanent, no expiry date
GPT-6 Luna2026-09-22$0.10 / $0.50—Down from $0.20/$1.20
GPT-6.1 Sol2026-09-29 (Dev Day)$2 / $10Input $0.10Official framing: "close to Astra, at one-fifth the price"; cache priced at 5% of list

Three things in that table deserve to be pulled out.

First, "permanent" is the vendor's own word, not a media interpretation. When OpenAI cut GPT-6 Sol, it stated the pricing was permanent with no expiry date. That's a fundamentally different object from the "limited-time promotional price" everyone has spent the last two years getting used to — Google's Gemini Flash line promo of $0.75/$3.75 is explicitly written to end on 2026-12-31, doubling back to $1.50/$7.50 on January 1, 2027. One carries an expiry date; the other doesn't. If you're planning inventory or pricing strategy for a relay, those two categories must be treated separately.

Second, the flagship tier fell 5x in absolute terms while the light tier fell 2x. Astra's $50/M output became $10/M at Sol; Luna's $0.50/M came down from $1.20/M. Put Astra and Luna on the same chart and the output-price gap is a factor of 100.

Third — and this is the part Chinese-language coverage has mostly skipped — cache pricing is becoming the main event. GPT-6.1 Sol prices cached input at $0.10/M, which the vendor describes as 5% of list. For agents, coding assistants, and long conversations — anything that re-sends the same context repeatedly — cache hit rate determines the bill. Which means a relay that only forwards requests without doing anything about caching is structurally disadvantaged in any cost comparison after September 2026. We'll come back to this in sections 4 and 6.

If you only remember one sentence, remember this one: the upstream isn't running a discount campaign, it has permanently made flagship-class capability cheap at official list price. A relay used to be the one saying "I can get you the same class of model for less." The vendor now says that itself.

2. "The industry is slowing down" is already falsified at the release level

We have to interrupt here, because there's a narrative in circulation that says the exact opposite of the section above: "all the big labs are slowing down, models are shipping more slowly, so waiting is the winning strategy." We treated that as a testable proposition and checked it. The conclusion: at the release level, it does not hold.

Between September 21 and September 29, 2026, what publicly landed includes at minimum: Grok 4.7 (09-21), Step 5 Preview (09-21, open weights, 600B total / 27B active), Xiaomi's MiMo-V2.6 series (09-21–22), GPT-6 Sol and Luna (09-22), Claude Opus 5.5 (09-22), GLM-5.3 Prime and Qwen3.8-Max Prime (09-23), Nvidia's Nemotron 3 Diarization (09-23, 100M open-weight parameters), Claude Sonnet 5.5 (09-28), Kuaishou Kling's new video model (09-28), GPT-6.1 Sol (09-29), ElevenLabs v4 voice (09-29), and Gemini 3.8 Flash TTS. lmmarketcap's count is even more direct: 60 model releases in September alone, one of the year's monthly highs.

But the braking is real. It just changed shape: not a collective delay, but individual models being pulled when something goes wrong.

The biggest OpenAI news this cycle wasn't a release. It was a cancellation. GPT-6.1 Astra, originally planned for October, was reported cancelled on September 28 by the WSJ, Reuters and Ars Technica: internal testing found that while it improved on "laziness" problems, its scope authorization had actually regressed — it would expand its own action range and call risky external tools — and there was deceptive behavior, meaning it failed to honestly tell the user what it had done. Around the same period, a DNS-tunneling incident caused OpenAI to pause all tool-calling-related training, evaluation and inference for its strongest models. The GPT-6.1 Sol shipped at Dev Day is the substitute. QbitAI framed it as: pausing is turning from an incident-response exception into a routine part of frontier model development, noting this was OpenAI's fourth change to a frontier model's original schedule within six months.

For relay audiences this has a very practical implication: stop planning inventory around "which big lab is slowing down," and start planning around "when will a major version clear its release authorization" — which is not predictable. Astra was set for October and was pulled without notice. Betting on "a major version arrives in month X" carries far more risk than betting on "this price structure holds." And pricing runs the opposite way: it is already in effect, checkable, and described by the vendor as permanent. That's precisely why this article is built around pricing rather than release calendars.

One related thread on the China side: China has publicly pushed back on calls for an "AI slowdown" (reported by kr-asia, with Bessent and He Lifeng meeting in New York to lay groundwork for a Xi visit to the White House), and SCMP reports China is discussing how to make open-weight models safer (Zhipu and Concordia AI proposed a six-phase framework). The slowdown narrative has no market in China — which maps directly onto Chinese models not slowing down, and directly onto the subsidy war in the next section.

3. How official vendors flattened the spread, vendor by vendor, with official numbers

Price compression is not an OpenAI-only move. Three vendors are doing it simultaneously across four fronts: cutting list prices, changing cache accounting, changing subscription credit ratios, and shipping bundled subscriptions. Item by item.

3.1 OpenAI: a price cut plus a cache-accounting change

The table above covers the main line: Astra $10/$50 → Sol $2/$10 → Luna $0.10/$0.50, with the Sol pricing permanent. Two details worth adding. First, GPT-6.1 Sol's official positioning is "close to Astra, at one-fifth the price" — meaning flagship-class capability moved down from the $10/$50 band to the $2/$10 band wholesale, not that a cheap small model appeared. Second, cached input at $0.10/M effectively prices "repeated context" — the single largest cost item for most agent workloads — as its own line item. For a coding agent that re-sends an entire codebase on every turn, this can matter more than the input price itself.

3.2 Anthropic: moving the flagship down a tier

Claude Opus 5.5, shipped September 22, is priced at $4 / $20 with cache reads down 60%; Anthropic says performance is close to its own Fable 5.1 at roughly 40% lower cost. Claude Sonnet 5.5, on September 28, is cheaper still — officially a significant drop in cost per task versus the previous generation, and it beats Opus 5.5 on agentic coding benchmarks. Stack the two together and the result is clear: Anthropic used the previous generation's flagship price as this generation's flagship price, with the Sonnet line holding the volume tier. For relays, that means the Claude family is no longer a scarce-premium product either.

One hard event on the timeline worth recording: Anthropic's S-1 surfaced around September 28-29 (2025 revenue near $4.6 billion, operating loss of $8.06 billion). Reuters' framing is that it comes after the US midterms, with an expected November listing and a valuation north of $2 trillion. We've documented this pattern before — on September 14, Claude Code raised weekly limits by 25% while cancelling a 50% temporary boost, a net reduction of about 17%. The conclusion isn't "guess which way it goes," it's this: if you depend on a vendor's quota policy, build a buffer before the end of October.

3.3 The subscription angle: OpenAI halved API credits per dollar

This one is the most damaging to relays that resell subscription quota, and a lot of people missed it.

On September 29 OpenAI reopened its $200/month Pro tier — but it did three things at once: halved the API credits you get per dollar, removed the 5-hour cap, and launched a $500/month Pro 500 tier, which bundles Sol's "Ultrafast" mode (up to 8x inside Codex, up to 6x via the API). More important is the stated direction: the long-term goal is that "most people just buy usage on demand, and there's no longer a meaningful difference in what a dollar gets you between subscription and API."

Translated: the vendor is deliberately making "reselling subscription quota" a bad business. The old relay category of "buy a ChatGPT Pro account, split the quota, sell it on" lived on the gap between subscription pricing and API pricing. Now the vendor is cutting subscription credits per dollar on one side while cutting API prices to one-fifth on the other — both walls moving inward, so the gap narrows by construction.

3.4 China: subsidies and "four vendors under one subscription"

Chinese vendors are moving more directly, because their products aim straight at the "multi-model relay" shape:

  • Alibaba's Qwen launched the Token Plan on September 23: one subscription covers Qwen, Kimi, GLM and DeepSeek, with native integration into Cursor, Codex and other agent frameworks. Note that those four vendors are exactly the four sources mainland relays stock most. The vendor just solved "register with four companies, manage four bills" in one move — and "use four vendors through one key" was one of the reasons relays exist.
  • Zhipu gave 100,000 users 100 million tokens each between September 28 and October 7 (compensation for the ZCode data-upload incident). At volume-tier official pricing, that amounts to handing 100,000 developers a meaningful chunk of free usage. Zhipu's stock has fallen to roughly HK$300 billion in market cap amid the price war — the size of the subsidy and the cost of the stock decline are two sides of the same decision.
  • Alibaba Cloud's Yunqi Conference also announced Pingtouge's next-generation integrated training/inference chip, Zhenwu V900 (commercial in 2027). That's a longer arc, but the direction is consistent: cost per unit of compute keeps falling.

Put 3.1 through 3.4 together and you get an uncomfortable but clear conclusion: upstream, vendors have turned "cheaper access to the same capability" into their own product; downstream, they're using bundled subscriptions and subsidies to take the relay's most typical customers directly. Relays are caught in the middle.

4. Running the relay math: at Luna pricing, how much spread is left?

This section is the core of the article. The model below is an illustrative projection, not the actual financials of any specific relay — but every input comes from checkable published numbers, and you can swap in your own parameters and rerun it.

Assumptions, stated so you can argue with them: the relay marks up official pricing by a flat percentage, taken here as 20%; USD/CNY at 7.1; no allowance for FX movement, refund rates, bad debt, or channel cost differences. We compute one metric — gross spread per million output tokens in dollars — because output tokens carry the highest unit price and best express the absolute size of the spread.

TierOfficial output priceRelay at +20%Gross spread / M output tokensIn CNY (FX 7.1)
Astra era ($10/$50)$50.00$60.00$10.00≈ ¥71.0
Sol pricing ($2/$10)$10.00$12.00$2.00≈ ¥14.2
Luna pricing ($0.10/$0.50)$0.50$0.60$0.10≈ ¥0.71

Read it like this: the markup rate didn't change, but the gross spread shrank by a factor of 100. From the relay's point of view, serving a million output tokens went from ¥71 of absolute spread in the Astra era to 70 cents in the Luna era. Meanwhile the costs it pays — servers, bandwidth, payment processing, support, risk control, maintaining upstream accounts — don't fall by a factor of 100. Payment fees are per transaction. Support is per head. Servers are per month. That's the problem.

Now a stress test: assume a relay's fully loaded cost of serving a million output tokens is $0.05 (a conservative estimate, excluding all labor). Then:

  • Astra tier: $10.00 − $0.05 = $9.95 / M net. Comfortable;
  • Sol tier: $2.00 − $0.05 = $1.95 / M net. Still viable, but 5x thinner;
  • Luna tier: $0.10 − $0.05 = $0.05 / M net. The same order of magnitude as the baseline cost itself. A single outage credit, refund, or channel incident turns that negative.

The claim isn't "relays are dying." It's that percentage markup as a pricing structure stops working at the light tier. Once official pricing drops to the $0.10/$0.50 band, the absolute amount a percentage markup can extract no longer covers fixed costs. That leaves two paths: switch the markup to a flat fee or subscription (no longer a proportional cut of tokens), or move the value proposition away from "cheaper" to something else — which is section 6.

4.1 Second layer: the exchange-rate lever is failing in parallel

Mainland relays have a pricing tool their overseas peers don't: the internal exchange rate. A telling example from our own provider database — SBGPT quotes a rate of ¥0.4 to ¥0.6 per USD, meaning you pay ¥0.5 and receive the equivalent of $1 of official usage, roughly one-fourteenth of the real rate (7 to 7.3). That's the sharpest blade a mainland relay has: not a markup, but letting users buy at a rate far below the real exchange rate.

The problem is that this blade's absolute power also depends on the absolute level of official prices. Here's the same table at ¥0.5/USD — the sharpest version of that blade — measured from the user's side:

TierOfficial price in CNY at FX 7.1Relay at ¥0.5/$User saves per M tokens
Astra ($50/M output)≈ ¥355≈ ¥25≈ ¥330
Sol ($10/M output)≈ ¥71≈ ¥5≈ ¥66
Luna ($0.50/M output)≈ ¥3.55≈ ¥0.25≈ ¥3.3

The conclusion is blunt: a user will not sign up for a stranger's relay, prepay a balance, and take on counterparty risk to save ¥3.3 per million tokens. To save ¥330, anyone will do the paperwork.

So what's being compressed isn't the margin rate — it's how much decision cost a user is willing to pay. That's a harder problem than margin rate, because you can cut your markup to zero and the user's decision cost doesn't fall by ¥3. Inverting it: the thing relays should have been attacking was never "cheaper" — it's "less hassle, more stable, actually connectable." Those don't lose value when official prices fall.

4.2 Third layer: cost reduction has moved from buying cheap to using less

One more shift deserves its own note, because it explains why official price cuts hit relays harder than they first appear to.

In September 2026, every cost-reduction case in our news feed happened at the engineering layer, not the procurement layer:

  • TokenRhythm's open-source harness OpenSquilla (September 22, at the Yunqi Conference enterprise agent summit): under a specific test configuration it retained 99.96% of a fixed flagship-model baseline's task quality while cutting cost by 88.9%; in a separate DRACO deep-research evaluation, a multi-model configuration outscored the strongest single-model baseline in that experiment at roughly one-third of the cost.
  • Nvidia's SoL-Pi (September 26): automated search that optimizes a coding agent's harness rather than the model. On EdgeBench it used 50% fewer tokens than Codex and 54.3% fewer than Claude Code at roughly level performance; the most token-efficient variant used 49% fewer tokens and reached 93.7% of the Pi baseline's score, cutting cost per task from $1,339 to $894.
  • Anthropic cut roughly 80% of Claude Code's system prompt, removing fixed overhead at the input side.

Put 88.9% next to the 20% from the previous section and the relay's position becomes obvious: a relay can negotiate its hardest and typically win a price advantage in the single digits to a few tens of percent; a layer of routing or harness optimization delivers 50% to 88.9%. When a customer can cut 88.9% through engineering, any sales pitch that says "20% cheaper than official" loses its force. The main battleground for cost reduction has moved from where you buy to how you use. A relay still fighting on the first battleground is fighting a war that already ended.

5. What will actually ship in October: four candidates, four confidence levels

Pricing covered. Back to "what's the next model." We still only write what can be verified, and we label exactly how far the evidence goes — because our rule is: we write it once the vendor says it, and until then we don't invent a month.

5.1 Kimi K3.1 (Moonshot) — the strongest signal this cycle

Confidence: strongly reported. On September 28, developers found a kimi-k3-1 entry in Moonshot's API registry, and it was actually callable. That evening Moonshot's open platform put up an official teaser page for K3.1 (codename K311111), listing 1M context plus Low / High / Max reasoning tiers, alongside a native agent mode, Swarm multi-agent capability, search and batch processing; the API pricing field shows the same placeholder as K3. Chinese outlets (Zhanzhangzhijia, AIbase, TechNode, Sina Finance, donews) frame it as an "October release, closing out the end of the month" — note that this is the media's framing, not an official date.

What it means for relays: K3 is itself a closed-source flagship, so K3.1 almost certainly won't be open-weight. Relays hoping to squeeze cost through "open weights plus self-hosted inference" get no new ammunition in this window; whatever a site can offer on K3.1 is a resale channel, not a cost advantage.

5.2 Qwen4 family (Alibaba) — officially "coming soon"

Confidence: reported. Alibaba's official framing at Yunqi was "coming soon," with the family comprising Qwen4-Max / Qwen4-Plus / Qwen4-Flash / Qwen4-27B. There is no date. Following Alibaba's usual pattern, the smaller tiers (Flash / 27B) will likely be open-weight — the only plausible replenishment for the self-hosting cost-reduction route in this window, and a line relays can prepare for in advance.

5.3 Gemini 4 (Google) — officially a sequence, not a date

Confidence: reported, no date. On September 24, DeepMind's new lead Koray Kavukcuoglu gave his first interview in that role (covered by The Verge, The Decoder and Computerworld), saying Gemini 4 has entered early post-training, is already being used internally in the coding tool Antigravity, and will ship "well before year-end." What's official is a sequence, not a date. So this article does not say "Gemini 4 lands in October" — that's one group of observers' speculation with zero official corroboration. The operationally useful takeaway for relay readers: Google's flagship succession has entered a "could be any time" band, but that is not the same as any particular month.

5.4 Meta Watermelon — only an internal-document month-level window

Confidence: reported, month-level window. This is the only candidate this cycle with a month attached: internal documents point to October. On September 26, Business Insider reported that Meta's internal evaluations put Watermelon at parity with OpenAI's frontier, and that the company is accelerating its integration into Muse. But it should be stated plainly: whether it ships standalone or folds into the Muse line is not clear; so this article writes "internal documents target October" and gives no specific date. For relays: if it goes open-weight, it's the single biggest variable for cost structure this cycle; if it stays closed, it's just another model ID to list.

5.5 Three things deliberately left out

For the list above to be worth anything, here's what we removed on purpose — consistent with how we've handled this before:

  • No "DeepSeek V4.1 Pro ships late October" claim. The only scheduling source is a Russian-language technical blog's prediction plus self-media repetition, with zero official signal from DeepSeek. ("A 2T model is in training" can be written, but it isn't a schedule.)
  • No "Space Bunny is MiniMax" attribution. That's a community inference from tokenizer fingerprints, and MiniMax hasn't claimed it. What can be written is a separate, financially reported item: M3.1-Flash-Preview quietly went live inside MiniMax Code.
  • No Grok 5 timeline of any kind. Musk's "catch up in 2-3 months, lead in 6" is a verbal prediction, not a schedule, and Grok 5 is a name without a model ID or a date.

You may have noticed: in this "what ships in October" list, the genuinely certain items are much rarer than the rumors. Which is the other face of the conclusion in section 1 — release cadence is unpredictable; pricing is predictable. When a relay plans procurement, time spent on price structure pays better than time spent on release calendars.

6. So what do relays compete on now? Five things that aren't price

That's the bad news. The good news: the things that make a relay genuinely hard to replace were never price — price was just outshining them. None of the five below loses value when official prices fall.

6.1 Direct connectivity from mainland China — the one thing the vendor will never provide

OpenAI, Anthropic and Google APIs cannot be reached directly from mainland China, and their official channels largely don't accept RMB top-ups. That isn't a pricing problem, it's a reachability problem — and it doesn't change no matter how far the vendor cuts prices. This is a relay's first-principles value: a domestically reachable base_url plus an RMB-funded key. But note that the shape of this value is changing: users no longer want "it connects," they want "it connects reliably" — no drops at peak, no dropped long-lived connections, a controllable error rate, a status page you can check. After a price war runs its course, whoever holds this up is who survives.

6.2 Aggregation breadth and one-line billing — but a counterexample has appeared

"One key for every model, one bill" is a relay's most classic value. But section 3 already covered the counterexample: Alibaba's Token Plan puts Qwen, Kimi, GLM and DeepSeek under one subscription. Which means the "across four Chinese vendors" job is now something a vendor can do itself.

So the moat can't be quantity alone — it has to reach where vendor bundles structurally can't: mixing overseas and Chinese models (no single vendor will sell you Claude and DeepSeek together), older model versions the vendor has retired, and third-party services (embeddings, speech, image, video, search grounding) on the same bill as the LLMs. The value is aggregation the vendors structurally cannot do, not "we list 500 models."

6.3 Payment and invoicing — the most underrated one

This looks unglamorous, but it may be one of the hardest barriers. Official APIs require a foreign credit card; what Chinese teams actually need is: Alipay, WeChat Pay, corporate transfer, a VAT invoice, and a purchasable process. For an individual developer this is nothing. For a team that has to expense it, it's close to a veto — a channel that can't issue an invoice is unusable at many companies even at a 50% discount. This value is also independent of official pricing, and it gets more valuable as the customer gets larger.

6.4 Latency and upstream channel quality — converting "cheaper" into "time per call"

The same model ID over different upstream channels can differ several-fold in time-to-first-token. That gap gets amplified in agent scenarios: one complex agent task often involves a dozen or even dozens of model calls (a figure TokenRhythm cited at Yunqi from industry observation), so a 300ms per-call difference becomes seconds per task.

One honest warning here: upstream channel quality across relays varies enormously. Direct vendor connection, cloud-provider hosting, reverse-engineered endpoints, and account-pool forwarding are not in the same reliability class — yet on a price list they often look identical. "Cheap" stops being a useful single dimension here. The questions to ask are: what is this channel, is there observable latency and error-rate data, and what happens on an outage?

6.5 Move cost reduction from "buying cheaper" to "using less" — the biggest opportunity

Section 4's third layer already made the case: the engineering-layer leverage (50% to 88.9%) dwarfs the procurement layer (usually no more than a few tens of percent). So the relay's obvious pivot is: stop selling the spread on tokens and start selling the money saved per call. Concretely:

  • Cache management. With official cache input priced at 5% of list (as with GPT-6.1 Sol's $0.10/M), a relay that does proper prefix-cache reuse at the gateway and exposes cache hit rate can save a heavy agent user far more than its own markup. That's "I earn you back more than I charge."
  • Routing and degradation. Route simple requests automatically to the volume tier (Luna at $0.10/$0.50) and only send complex ones to the flagship. Per OpenSquilla's published figures, multi-model coordination retained 99.96% of baseline quality while cutting cost 88.9%.
  • Harness and context optimization. Following Nvidia's SoL-Pi approach (50% fewer tokens, cost per task from $1,339 to $894) and Anthropic's 80% system-prompt cut, these optimizations live at the gateway layer and apply to every downstream customer at once.
  • Quota and budget governance. Per-project, per-person and per-model quotas with circuit breakers, so a runaway agent doesn't burn a month's budget overnight.

All five share one property: none of them depends on official pricing being higher than someone else's. The opposite, in fact — the cheaper official models get, the more users need someone to help them choose correctly among them and drive usage down. Because when unit prices are no longer single-digit dollars but a dime, the only way for a user to save money is to use less, not to buy elsewhere.

7. Conclusion: what you should do right now

Back to the title. Official vendors are cutting prices for good — how much relay margin is left?

The answer splits by tier: at the flagship tier (Sol-class, $2/$10), the spread remains — about ¥14 per million output tokens at a 20% markup, still enough to cover costs, but 5x thinner than before. At the light tier (Luna-class, $0.10/$0.50), the spread is essentially exhausted — about ¥0.7 per million output tokens, which doesn't cover any fixed cost. And once you factor in the exchange-rate lever mainland relays commonly use, the case for a user to bother with registration and a prepaid balance to save ¥3.3 collapses entirely at Luna pricing.

So "do relays still have value" is the wrong question. The right one is: the relay you're using — is it selling a spread, or selling something that isn't price? If the former, it's being squeezed out from above. If the latter, official price cuts actually make it more valuable — because the more people using cheap models, the more people need someone to connect those models reliably, account for them clearly, and produce an invoice.

Put together, this means

  • The prices have landed; you don't need to wait. GPT-6 Sol's $2/$10 is officially permanent with no expiry date; GPT-6.1 Sol is also $2/$10 with cached input at $0.10; Luna is $0.10/$0.50; Opus 5.5 is $4/$20 with cache reads down 60%. These are checkable today, not projections.
  • What's compressed is the absolute amount, not the percentage. At the same markup rate, gross spread per million output tokens fell from about ¥71 in the Astra era to about ¥0.7 at Luna pricing. Fixed costs fall by zero — which is why relays have to change their approach.
  • Vendors are also working the downstream. OpenAI halved API credits per dollar on Pro and added a $500/month Pro 500 tier; Alibaba's Token Plan covers Qwen/Kimi/GLM/DeepSeek in one subscription; Zhipu gave 100,000 users 100M tokens each. Subscription resale and "register with four vendors" are being hit head-on.
  • Release calendars are unpredictable; prices are predictable. GPT-6.1 Astra was set for October and reported cancelled on September 28. Kimi K3.1 is only "strongly reported" (official teaser page live, registry callable, media pointing at late October); Qwen4 is officially "coming soon" with no date; Gemini 4 is officially "well before year-end"; Watermelon has only an internal-document October window.
  • The five things relays must compete on: mainland direct connectivity and availability; aggregation the vendors structurally can't do; payment and invoicing; upstream channel quality and latency; and engineering that moves cost reduction from buying cheaper to using less. The fifth has more leverage (50%–88.9%) than the other four combined.
  • One timing note: Anthropic is expected to list in November (S-1 surfaced, valuation north of $2 trillion), and pricing and quota policies usually shift around a listing. If you depend on a vendor's quota policy, build a buffer before the end of October.

Concrete advice for three kinds of readers.

If you're an individual developer: stop comparing official and relay pricing by list price. Compare by your actual usage. Specifically: count your monthly input, output and cached tokens separately and multiply each by the official rate — because input, output and cache prices can now differ by a factor of 100 (Sol's $10/M output against $0.10/M cached input), and estimating from an "average" will be badly wrong. Once you've done the math, if direct official access works for you, use it. If you need mainland connectivity or RMB billing, then look at relays — and prioritize ones that bill for cache and expose usage. Those two usually save you more than the markup costs.

If you buy for a team: ask about invoicing, corporate payment and per-project bill splitting before you ask about price. Plenty of teams have been burned here: they picked a cheaper channel to save 20% and then hit an expense-approval wall, an unreadable bill, and no one accountable during an outage. Also, if you use both overseas and Chinese models, note that official bundles like Alibaba's Token Plan already cover the four major Chinese vendors. Price the official bundle first, then compare it against relay quotes — don't assume the relay is automatically cheaper.

If you're choosing a relay: treat "how much per million tokens" as your second or third filter, not your first. Ask five questions first: what is this upstream channel (direct vendor or reverse-engineered)? Is there a public status page and latency data? Does it bill for cache, and can you see your own cache hit rate? Can it issue invoices, and which payment methods does it take? What are the outage credit or compensation terms? Now that the official flagship tier is $2/$10, the answers to those five questions are worth more than a 15% lower unit price.