SiliconFlow Review: Pricing & Comparison
Domestic open-source model cloud hosting — DeepSeek, Qwen, GLM and 100+ models, an in-house inference engine plus domestic chip support
Last verified: 2026-08-11 · Visit official site →
SiliconFlow isn’t a relay
This distinction matters.
A typical relay’s business model: bulk-purchase official OpenAI/Anthropic accounts, then resell access at a markdown to the official exchange rate. The platform itself doesn’t hold any model weights.
SiliconFlow’s business model is entirely different: it hosts open-source model weights directly, running its own inference engine (already adapted to domestic Ascend chips), making it one of the actual cloud operators of models like DeepSeek, Qwen, and Kimi K2.
That means: no “capability downgrade” risk (there’s no such thing as official-account rate-limiting here); no extra markup from an intermediary layer; and model version updates track official open-source releases almost in sync. The trade-off: no GPT/Claude/Gemini or other closed-source commercial models.
Pricing in detail
SiliconFlow’s pricing sits at the floor of the market among comparable platforms — some small-parameter models are permanently free, and even its flagship models are priced extremely low. Per August 2026 site data, reference pricing for major models (check the live site for current pricing):
| Model | Input price (per M tokens) | Output price (per M tokens) | Direct connect |
|---|---|---|---|
| DeepSeek-V4-Flash | ¥1.00 | ¥2.00 (cache ¥0.02) | ✓ |
| DeepSeek-V4-Pro | ¥12.00 | ¥24.00 (cache ¥1.00) | ✓ |
| Qwen3.5-397B-A17B | ¥1.20 | ¥7.20 | ✓ |
| Qwen2.5-7B (free) | ¥0 | ¥0 | ✓ |
| Kimi-K2.7-Code | ¥6.50 | ¥27.00 (cache ¥1.30) | ✓ |
| GLM-5.2 | ¥8.00 | ¥28.00 (cache ¥2.00) | ✓ |
| MiniMax-M2.5 | ¥2.10 | ¥8.40 (cache ¥0.21) | ✓ |
Take DeepSeek-V4-Flash as an example: roughly ¥1.00 input and ¥2.00 output per million tokens — still among the lowest on comparable platforms. For comparison, Anthropic’s official Claude Sonnet 4.6 pricing is $3/$15 per million tokens (input/output), roughly ¥21.6/¥108 — SiliconFlow’s flagship open-source models are priced one to two orders of magnitude below closed-source flagships.
Worth noting: SiliconFlow doesn’t offer GPT/Claude/Gemini or other closed-source commercial models. If your use case requires Claude Code or GPT-5.5, you’ll need a different relay.
Model coverage and real-world performance
SiliconFlow currently lists 150+ open-source models, including:
- The DeepSeek family: DeepSeek-V3.2, DeepSeek-V4-Flash, DeepSeek-V4-Pro, the DeepSeek-R1 series, and more — updated almost in sync with official releases.
- The Qwen family: Qwen3-235B-A22B, Qwen3-30B-A3B, Qwen2.5-7B (free), and the rest of the Qwen lineup.
- GLM/InternLM: Zhipu’s GLM series, Shanghai AI Lab’s InternLM series.
- Kimi K2 Instruct: Moonshot’s Kimi K2 model, newly listed in 2026.
- MiniMax M3: MiniMax’s latest flagship model.
Since SiliconFlow focuses on open-source models, it doesn’t have GPT-5.5, Claude Fable 5, Gemini 3.5, or other closed-source flagships. Some reviews note that SiliconFlow’s support for domestic chips (Ascend) is solid, with inference performance already on par with mainstream GPUs. There’s no capability-downgrade risk, since it directly hosts open-source weights rather than depending on an official account subject to rate limits.
Using it day to day
Signup: go to siliconflow.cn, register with an email address to start using the service; after completing real-name verification, claim a ¥16 platform-wide universal voucher via Activity Center → 认证专享礼. Note: since 2026-05-15, unverified accounts can’t use the platform. The whole process takes 3-5 minutes.
Payment: supports Alipay and WeChat Pay, settled in RMB, with no minimum top-up (though we’d suggest topping up at least ¥10 for full feature access).
Getting your API key: log in and generate one directly on the “API Keys” page in the console. It supports OpenAI-compatible format (https://api.siliconflow.cn/v1) — you can go from signup to your first API call within 5 minutes.
Dashboard: gives usage statistics, cost breakdowns, and call logs — relatively lean, functional but not as detailed as some enterprise-grade platforms.
Referral program: a public “Referral Officer” two-sided bonus program (campaign 2026-01-15 through 2026-12-31): become a Referral Officer and each successful referral (signup + real-name verification) earns both sides a ¥16 universal voucher, no business-development contact needed; the referral link is available directly in the console.
Stability and latency
SiliconFlow’s servers are deployed domestically, no proxy/VPN needed — friendly to mainland developers. Based on user-reported real-world results, generating 500 words of content takes roughly 7 seconds (for the DeepSeek-V3 family), saving the cross-border network time compared to an official API.
There’s no official unified latency figure or SLA commitment, but given the mainland direct-connect advantage, real-world latency is typically better than an overseas relay requiring a proxy. No large-scale outages have been observed so far, and individual-developer feedback on stability has been positive.
On rate limits: free models have a QPS cap (check the live site for the current configuration); paid models have relatively looser rate limits — contact support directly for high-concurrency scenarios.
Who it’s for
- Heavy users of domestic open-source models: the primary use case for DeepSeek, Qwen, Kimi, and similar models — far cheaper than any closed-source relay.
- Vibe-coding individual developers: validate product prototypes cheaply — ¥14 in trial credit goes a long way for testing.
- Anything needing mainland direct connect: projects running on mainland servers that don’t want to configure a separate proxy just for API calls.
- Teams that want a backup model pool: using open-source models as a fallback below Claude/GPT, with a significant cost advantage.
Not a good fit: use cases that need closed-source flagships like GPT-5.5, Claude Code, or Gemini.
Head-to-head comparisons
Compared to similar domestic platforms:
- vs. 147API: 147API covers closed-source models like OpenAI/Claude/Gemini, but its pricing generally runs higher than SiliconFlow’s open-source model pricing; if you only use open-source models, SiliconFlow is cheaper.
- vs. 302.AI: 302.AI supports closed-source flagships (GPT/Claude/Gemini) with a balance that never expires, priced close to official rates; if you need one-stop access to multiple closed-source models, 302.AI is the better fit.
Core distinction: SiliconFlow = floor-price access to open-source models, while 302.AI/147API = a convenient channel to closed-source models.
Bottom line
SiliconFlow is one of the go-to platforms for domestic open-source large-model APIs — highly competitive pricing (some models free), mainland direct connect with no proxy needed, and a friendly signup experience. Its one real limitation is not covering GPT/Claude or other closed-source flagships — if your use case requires a closed-source model, you’ll need to pair it with another relay.
For developers whose primary workload is DeepSeek, Qwen, and other domestic open-source models, we’d strongly recommend it as a primary platform.
If you also need Claude and want to run integration tests on free credit, AiHubMix’s permanently free test tier plus Prompt Caching support can supplement SiliconFlow’s closed-source-side gap. If even the low-price tier feels too expensive and you want to push API exploration cost as low as possible, Yunwu API (0.5¥/USD plus a daily free GitHub credit) is a more aggressive low-price alternative.
Trial credit and promotion details are subject to SiliconFlow’s live site. This review was verified 2026-08-11.
Related reviews
- AICloud Feiyun: 50 free Sonnet calls daily, low-cost direct-connect Claude access
- Boxying: a lightweight, no-frills domestic relay, mainstream models, quick to get started
- Lumin AI: launched in 2026, low ¥5 entry bar, Kiro endpoint at ¥2/10M tokens
- DMXAPI: 480+ models, multimodal aggregation across text/image/video, enterprise-grade stability
Quick facts
| Pricing model | Pay-as-you-go, claims the lowest prices in the market; after real-name verification claim one ¥16 platform-wide universal voucher via Activity Center → 认证专享礼; the "Referral Officer" program pays both sides ¥16 when an invited friend completes signup + real-name verification (campaign through 2026-12-31) |
|---|---|
| Model coverage | 100+ open-source models spanning DeepSeek, Qwen, GLM, InternLM and other mainstream domestic open-source large models; some models are free to call |
| Latency / SLA | Mainland servers with direct connect, no cross-border network overhead; no official public latency figures |
| Mainland direct connect | Mainland direct connect, no proxy/VPN needed |
| Best for | Developers / Fallback / backup |
| Referral program | A public referral program: new users get ¥14 in trial credit on signup, and referrers get another ¥14 for each successful referral — no business-development contact required, making it one of the few programs a cold-start project can plug into with zero negotiation. |
Pros
- One of the broadest open-source model line-ups among domestic platforms — the DeepSeek/Qwen series update quickly
- Mainland direct connect — lower latency and lower compliance risk than overseas relays
- A public two-sided referral program — usable as promotional material right out of the gate during a cold start, with no business-development process required
- Well suited to budget-constrained individual developers and vibe-coding projects doing prototype validation
Cons
- Focused on open-source models — coverage of closed-source flagships (official GPT-4.x/Claude 4.x) trails some relays
- No unified public latency/SLA figures — enterprise-grade guarantees aren't as strong as SLA-focused platforms like Shiyun API
Compare more AI API relays
See the full comparison board — filter by price tier, model coverage, and mainland direct-connect status.
Back to the comparison board →