OpenRouter Review: Pricing & Comparison
A cross-vendor multi-model aggregator — 400+ models across 70+ providers, one API key to call OpenAI/Anthropic/Google/Meta/Mistral and more
Last verified: 2026-08-11 · Visit official site →
OpenRouter isn’t a “relay” in the domestic Chinese sense
Mainland Chinese developers searching for a “relay” (中转站) usually have two specific needs in mind: Alipay/WeChat Pay settlement, and mainland direct connect for low latency. OpenRouter satisfies neither — payment is mostly credit-card based, and its nodes are overseas, so mainland access needs a proxy.
But OpenRouter offers something else that’s hard to replace: the broadest model coverage in the world — 400+ models across 70+ vendors, all reachable through a single API key spanning OpenAI, Anthropic, Google, Meta, Mistral, Cohere, and nearly every other major provider.
If your core need is comparing multiple models side by side, or you already have proxy infrastructure and don’t care about mainland direct connect, OpenRouter is the natural first choice; if you need low-latency mainland direct connect, look at other platforms in this review like 4SAPI or KoalaAPI.
Pricing: official pass-through + a 5.5% fee
OpenRouter’s core pricing logic is “pass through the official price + a platform fee.” Breaking that down:
Pass-through portion: per-token pricing for major models is essentially in line with each vendor’s official pricing (as of March 2026 data; re-verified 2026-08 — OpenAI’s homepage now promotes the GPT-5.6 family, so individual prices should be checked against the official models page):
| Model | Input price (per M tokens) | Output price (per M tokens) | Direct connect |
|---|---|---|---|
| Claude Fable 5 | $10 | $50 | ✗ (proxy required) |
| Claude Opus 4.8 | $5 | $25 | ✗ |
| Claude Sonnet 4.6 | $3 | $15 | ✗ |
| GPT-5.5 | $1.25 | $10 | ✗ |
| Gemini 3.5 Flash | see official site | see official site | ✗ |
| DeepSeek-V4-Flash | see official site | see official site | ✗ |
Platform fee (the hidden cost):
- Credit-card payment fee: 5.5% (minimum $0.80 — the effective rate is higher on small top-ups)
- A 5% BYOK fee once monthly call volume exceeds 1 million requests
Using Claude Sonnet 4.6 as an example: official pricing is $3/$15 per million tokens; the actual cost on OpenRouter comes to roughly $3.17/$15.83 (+5.5%). In practice, budget for official pricing plus a 5-7% surcharge as a safe planning assumption.
On the free side, there are currently 27+ fully free open-source models (input and output both $0), including Gemma 4 and MiniMax M2.5, capped at 50 calls a day (the limit loosens once you’ve topped up $10 or more).
Model coverage breadth: a global top tier
OpenRouter’s core strength is the breadth of its model coverage:
- OpenAI: full coverage of GPT-5.5, GPT-4o, and the o1/o3 series
- Anthropic: Claude Fable 5, Claude Opus 4.8, Claude Sonnet 4.6, Claude Haiku 4.5
- Google: Gemini 3.5 Flash/Pro, the Gemini 2.5 series
- Meta: the Llama 4 series
- Mistral: Mistral Large 3, Mistral Small, and others
- Others: Grok 4.3 (xAI), DeepSeek V4, Qwen3, Kimi K2, and more
New models tend to go live extremely fast — usually within hours to days of a vendor’s public release. On the automatic-failover side, users can configure “auto-switch to a backup model when the primary is unavailable,” giving production deployments a meaningful degree of fault tolerance.
Using it day to day
Signup: go to openrouter.ai, sign up with a Google or GitHub account in one click — new users get no signup credit; after signing up the default free tier applies (50 free calls/day across 25+ open-source models, rate-limited). The whole process takes about 2 minutes.
Payment: mainly supports international credit cards and USDC crypto; some pages support Alipay (via the “use one-time payment methods” option), but the flow is relatively involved. Mainland users without an overseas credit card face a somewhat higher barrier to entry.
Getting your API key: after signing up, generate a key on the Keys page. The endpoint is https://openrouter.ai/api/v1, fully OpenAI SDK compatible — just swap the base_url and api_key.
Dashboard: gives detailed call logs, cost breakdowns, and per-model/per-time-period statistics — good for cost analysis and multi-model comparison.
Support: mainly English-language community — Discord and GitHub are both active — but dedicated support channels for mainland users are weaker.
Stability and latency
OpenRouter has no mainland direct-connect nodes — servers are deployed overseas. Based on 2026 real-world data:
- Through a proxy: first-token latency around 600-1200ms
- Without a proxy (direct mainland access): latency around 1500-3000ms, and unstable
For latency-sensitive AI coding tools like Claude Code and Cursor, we’d recommend running a stable proxy the whole time, or timeouts become routine. Compared to 4SAPI (which claims a 520ms TTFT), OpenRouter’s experience on mainland networks is noticeably worse.
The failover mechanism is a real highlight: when an upstream model (say, GPT-5.5) gets rate-limited, it can automatically switch to a backup model, reducing business disruption.
Who it’s for
- Developers doing multi-model benchmarking/selection: if your vibe-coding project wants to compare Claude vs. GPT vs. Gemini on quality and cost, OpenRouter lets you test them all with one key.
- Teams that already have overseas proxy infrastructure: latency and payment aren’t obstacles in this scenario.
- Product teams building for an overseas audience: needing international model selection where the call chain runs overseas anyway.
- Developers who want to try free models: 27+ free open-source models let you validate prompts at zero cost.
Not a good fit: projects running primarily on mainland servers that need low-latency direct connect, or users limited to Alipay/WeChat Pay.
Head-to-head comparisons
- vs. 4SAPI: 4SAPI offers a mainland direct-connect node (claiming 520ms TTFT) and supports Alipay, giving a noticeably better domestic experience than OpenRouter. OpenRouter has broader model coverage and a more mature international footprint, but its domestic experience clearly trails 4SAPI.
- vs. 302.AI: 302.AI is also pay-as-you-go with a balance that never expires, and supports Alipay and mainland direct connect — better suited to mainland users who want a one-stop AI toolset, not just a bare API. OpenRouter leans more toward developer-layer model routing and multi-model comparison.
The core decision rule: if you’re in mainland China and mainly use Claude/GPT, choose 4SAPI or 302.AI; if you’re operating globally and need the broadest possible model coverage, choose OpenRouter.
Practical recommendations
As a primary platform (good fit): teams that already have stable proxy infrastructure; products serving an overseas audience; technical selection processes that need to benchmark 10+ candidate models side by side.
As a supplementary tool (good fit): mainland day-to-day developers can treat OpenRouter as a “multi-model benchmarking tool” — occasionally running a multi-model comparison, while keeping daily production work on a mainland direct-connect platform.
Not recommended as a primary platform (poor fit): businesses sensitive to mainland access latency; users limited to Alipay/WeChat Pay; developers who need low-latency Claude Code usage day-to-day (4SAPI or KoalaAPI are a better fit).
Markup percentages, free-credit terms, and other details are subject to OpenRouter’s live site and billing page. This review was verified 2026-08-11.
Related reviews
- Cloudflare AI Gateway: a free AI traffic-control layer — caching, rate limiting, monitoring, edge acceleration
- lxg2it ModelRouter: 7+ providers at 0% markup, auto-routes to the cheapest available model, OpenAI-compatible
- Zhipu AI: Zhipu’s official GLM series, integrated multimodal AI, top-tier Chinese-language capability
- NodAPI: fast multi-model aggregation, low-latency mainland direct connect, a top pick for real-time scenarios
Quick facts
| Pricing model | Passes through official pricing plus a markup (different sources cite inconsistent figures — 1%, 5.5%, up to 25%), includes 25+ free-tier models (rate-limited), and gives new users $1 in free credit |
|---|---|
| Model coverage | 400+ models across 70+ providers, spanning nearly every major vendor — OpenAI, Anthropic, Google, Meta, Mistral, Cohere, and more |
| Latency / SLA | Automatic failover to a backup model when the primary is unavailable; latency ultimately depends on the upstream official API |
| Mainland direct connect | Overseas nodes — mainland China access typically requires a proxy/VPN |
| Best for | Developers / Fallback / backup |
| Referral program | No public Chinese-market affiliate/referral program found; the common global pattern is a developer-community or integrator revenue share — unconfirmed whether it's open to mainland China users. |
Pros
- The broadest model coverage of any platform in this review — nearly every major vendor and newly released model gets integrated almost immediately
- Ideal for "multi-model benchmarking/selection" — one key lets you test every candidate model and compare cost and quality side by side
- Automatic failover — when an upstream model is rate-limited or down, it can switch smoothly to a backup
- 25+ free models available (rate-limited), good for zero-cost prompt testing
- Highly international — mature docs and community, well suited to developer tools aimed at a global audience
Cons
- Mainland China access is poor without a proxy/VPN, which somewhat defeats the point of a "relay" for domestic developers
- Markup percentage isn't consistently described across models or sources — you need to calculate your real cost yourself
- Payment is mainly international credit card/PayPal, a higher barrier for users who only have Alipay/WeChat Pay
Compare more AI API relays
See the full comparison board — filter by price tier, model coverage, and mainland direct-connect status.
Back to the comparison board →