Free / Budget Broad coverage (100+) Mainland direct connect, no proxy/VPN needed ★ 4.2 / 5

SiliconFlow Review: Pricing & Comparison

Domestic open-source model cloud hosting — DeepSeek, Qwen, GLM and 100+ models, an in-house inference engine plus domestic chip support

Last verified: 2026-08-11 · Visit official site →

SiliconFlow isn’t a relay

This distinction matters.

A typical relay’s business model: bulk-purchase official OpenAI/Anthropic accounts, then resell access at a markdown to the official exchange rate. The platform itself doesn’t hold any model weights.

SiliconFlow’s business model is entirely different: it hosts open-source model weights directly, running its own inference engine (already adapted to domestic Ascend chips), making it one of the actual cloud operators of models like DeepSeek, Qwen, and Kimi K2.

That means: no “capability downgrade” risk (there’s no such thing as official-account rate-limiting here); no extra markup from an intermediary layer; and model version updates track official open-source releases almost in sync. The trade-off: no GPT/Claude/Gemini or other closed-source commercial models.

Pricing in detail

SiliconFlow’s pricing sits at the floor of the market among comparable platforms — some small-parameter models are permanently free, and even its flagship models are priced extremely low. Per August 2026 site data, reference pricing for major models (check the live site for current pricing):

ModelInput price (per M tokens)Output price (per M tokens)Direct connect
DeepSeek-V4-Flash¥1.00¥2.00 (cache ¥0.02)
DeepSeek-V4-Pro¥12.00¥24.00 (cache ¥1.00)
Qwen3.5-397B-A17B¥1.20¥7.20
Qwen2.5-7B (free)¥0¥0
Kimi-K2.7-Code¥6.50¥27.00 (cache ¥1.30)
GLM-5.2¥8.00¥28.00 (cache ¥2.00)
MiniMax-M2.5¥2.10¥8.40 (cache ¥0.21)

Take DeepSeek-V4-Flash as an example: roughly ¥1.00 input and ¥2.00 output per million tokens — still among the lowest on comparable platforms. For comparison, Anthropic’s official Claude Sonnet 4.6 pricing is $3/$15 per million tokens (input/output), roughly ¥21.6/¥108 — SiliconFlow’s flagship open-source models are priced one to two orders of magnitude below closed-source flagships.

Worth noting: SiliconFlow doesn’t offer GPT/Claude/Gemini or other closed-source commercial models. If your use case requires Claude Code or GPT-5.5, you’ll need a different relay.

Model coverage and real-world performance

SiliconFlow currently lists 150+ open-source models, including:

  • The DeepSeek family: DeepSeek-V3.2, DeepSeek-V4-Flash, DeepSeek-V4-Pro, the DeepSeek-R1 series, and more — updated almost in sync with official releases.
  • The Qwen family: Qwen3-235B-A22B, Qwen3-30B-A3B, Qwen2.5-7B (free), and the rest of the Qwen lineup.
  • GLM/InternLM: Zhipu’s GLM series, Shanghai AI Lab’s InternLM series.
  • Kimi K2 Instruct: Moonshot’s Kimi K2 model, newly listed in 2026.
  • MiniMax M3: MiniMax’s latest flagship model.

Since SiliconFlow focuses on open-source models, it doesn’t have GPT-5.5, Claude Fable 5, Gemini 3.5, or other closed-source flagships. Some reviews note that SiliconFlow’s support for domestic chips (Ascend) is solid, with inference performance already on par with mainstream GPUs. There’s no capability-downgrade risk, since it directly hosts open-source weights rather than depending on an official account subject to rate limits.

Using it day to day

Signup: go to siliconflow.cn, register with an email address to start using the service; after completing real-name verification, claim a ¥16 platform-wide universal voucher via Activity Center → 认证专享礼. Note: since 2026-05-15, unverified accounts can’t use the platform. The whole process takes 3-5 minutes.

Payment: supports Alipay and WeChat Pay, settled in RMB, with no minimum top-up (though we’d suggest topping up at least ¥10 for full feature access).

Getting your API key: log in and generate one directly on the “API Keys” page in the console. It supports OpenAI-compatible format (https://api.siliconflow.cn/v1) — you can go from signup to your first API call within 5 minutes.

Dashboard: gives usage statistics, cost breakdowns, and call logs — relatively lean, functional but not as detailed as some enterprise-grade platforms.

Referral program: a public “Referral Officer” two-sided bonus program (campaign 2026-01-15 through 2026-12-31): become a Referral Officer and each successful referral (signup + real-name verification) earns both sides a ¥16 universal voucher, no business-development contact needed; the referral link is available directly in the console.

Stability and latency

SiliconFlow’s servers are deployed domestically, no proxy/VPN needed — friendly to mainland developers. Based on user-reported real-world results, generating 500 words of content takes roughly 7 seconds (for the DeepSeek-V3 family), saving the cross-border network time compared to an official API.

There’s no official unified latency figure or SLA commitment, but given the mainland direct-connect advantage, real-world latency is typically better than an overseas relay requiring a proxy. No large-scale outages have been observed so far, and individual-developer feedback on stability has been positive.

On rate limits: free models have a QPS cap (check the live site for the current configuration); paid models have relatively looser rate limits — contact support directly for high-concurrency scenarios.

Who it’s for

  • Heavy users of domestic open-source models: the primary use case for DeepSeek, Qwen, Kimi, and similar models — far cheaper than any closed-source relay.
  • Vibe-coding individual developers: validate product prototypes cheaply — ¥14 in trial credit goes a long way for testing.
  • Anything needing mainland direct connect: projects running on mainland servers that don’t want to configure a separate proxy just for API calls.
  • Teams that want a backup model pool: using open-source models as a fallback below Claude/GPT, with a significant cost advantage.

Not a good fit: use cases that need closed-source flagships like GPT-5.5, Claude Code, or Gemini.

Head-to-head comparisons

Compared to similar domestic platforms:

Core distinction: SiliconFlow = floor-price access to open-source models, while 302.AI/147API = a convenient channel to closed-source models.

Bottom line

SiliconFlow is one of the go-to platforms for domestic open-source large-model APIs — highly competitive pricing (some models free), mainland direct connect with no proxy needed, and a friendly signup experience. Its one real limitation is not covering GPT/Claude or other closed-source flagships — if your use case requires a closed-source model, you’ll need to pair it with another relay.

For developers whose primary workload is DeepSeek, Qwen, and other domestic open-source models, we’d strongly recommend it as a primary platform.

If you also need Claude and want to run integration tests on free credit, AiHubMix’s permanently free test tier plus Prompt Caching support can supplement SiliconFlow’s closed-source-side gap. If even the low-price tier feels too expensive and you want to push API exploration cost as low as possible, Yunwu API (0.5¥/USD plus a daily free GitHub credit) is a more aggressive low-price alternative.

Trial credit and promotion details are subject to SiliconFlow’s live site. This review was verified 2026-08-11.

  • AICloud Feiyun: 50 free Sonnet calls daily, low-cost direct-connect Claude access
  • Boxying: a lightweight, no-frills domestic relay, mainstream models, quick to get started
  • Lumin AI: launched in 2026, low ¥5 entry bar, Kiro endpoint at ¥2/10M tokens
  • DMXAPI: 480+ models, multimodal aggregation across text/image/video, enterprise-grade stability

Quick facts

Pricing modelPay-as-you-go, claims the lowest prices in the market; after real-name verification claim one ¥16 platform-wide universal voucher via Activity Center → 认证专享礼; the "Referral Officer" program pays both sides ¥16 when an invited friend completes signup + real-name verification (campaign through 2026-12-31)
Model coverage100+ open-source models spanning DeepSeek, Qwen, GLM, InternLM and other mainstream domestic open-source large models; some models are free to call
Latency / SLAMainland servers with direct connect, no cross-border network overhead; no official public latency figures
Mainland direct connectMainland direct connect, no proxy/VPN needed
Best forDevelopers / Fallback / backup
Referral programA public referral program: new users get ¥14 in trial credit on signup, and referrers get another ¥14 for each successful referral — no business-development contact required, making it one of the few programs a cold-start project can plug into with zero negotiation.

Pros

  • One of the broadest open-source model line-ups among domestic platforms — the DeepSeek/Qwen series update quickly
  • Mainland direct connect — lower latency and lower compliance risk than overseas relays
  • A public two-sided referral program — usable as promotional material right out of the gate during a cold start, with no business-development process required
  • Well suited to budget-constrained individual developers and vibe-coding projects doing prototype validation

Cons

  • Focused on open-source models — coverage of closed-source flagships (official GPT-4.x/Claude 4.x) trails some relays
  • No unified public latency/SLA figures — enterprise-grade guarantees aren't as strong as SLA-focused platforms like Shiyun API

Compare more AI API relays

See the full comparison board — filter by price tier, model coverage, and mainland direct-connect status.

Back to the comparison board →