If you only remember one number, make it this: on Artificial Analysis's Intelligence Index, Muse Spark 1.2 costs roughly $0.40 per task to reach its intelligence tier — versus $0.86 for Kimi K3 (max) and $1.18 for GPT-5.5 (xhigh) at the same rough intelligence level. Only Grok 4.5 (high, $0.37) and GPT-5.6 Sol (medium, $0.39) beat it on that metric. This is the third model MSL has shipped in just four months, so let's be clear about what it is first: Muse Spark is not an image or video generation model. It's a large language model built for multimodal reasoning and coding-agent workflows, and this particular 1.2 update is squarely about coding — it launched alongside Muse Code, Meta's first terminal-based coding agent.
This piece works through, in order: what kind of model Muse Spark actually is (clearing up the "is this an image generator" confusion first); the full version history from the April 8 launch of 1.0 through the August 5 release of 1.2; four genuine highlights worth digging into (cost-efficiency, Muse Code's persistent background-agent architecture, the pace of the benchmark curve itself, and native multimodality plus a 1M-token context window); the limitations that need to be said out loud (the gap between self-reported and independently verified scores, closed weights with no self-hosting, Muse Code's missing platform/IDE support, and the privacy trade-off of the cheaper "contributor" pricing tier); a side-by-side comparison table; and — the question that matters most to our readers — what this means for mainland China AI API relay access right now. Every number here is cross-checked against Meta's own research blog, Artificial Analysis, Vals AI, OpenRouter, VentureBeat, MarkTechPost, TechCrunch, and TechTimes. Wherever sources disagree, or a claim lacks third-party verification, we say so explicitly rather than smoothing it over to make the review look tidier than the facts allow.
Contents
- 1. What Muse Spark is: clearing up a common misconception first
- 2. Version history: from April's "rebuild from scratch" to the August 5 1.2 release
- 3. Highlight one: cost-efficiency — one of the best "intelligence per dollar" options at its tier
- 4. Highlight two: Muse Code's architecture — persistent background agents plus isolated git-worktree sub-agents
- 5. Highlight three: the benchmark curve itself — two tiers in three months, closing in on Claude Opus 5
- 6. Highlight four: native multimodality, a 1M-token window, and it's already usable today
- 7. The cold water that has to be thrown: how big is the self-reported vs. verified gap
- 8. Other real limitations: closed weights, an immature Muse Code, and the privacy cost of the cheap tier
- 9. Side by side: Muse Spark 1.2 against the current crop of leading models
- 10. What this means for mainland China AI API relay readers
- 11. Who should use it now, and who should wait
- 12. Conclusion
1. What Muse Spark is: clearing up a common misconception first
Let's settle the most basic question first, because the name "Muse" naturally invites comparisons to image or video generation tools — but that's not what this is. Muse Spark is a multimodal reasoning large language model built by Meta Superintelligence Labs (MSL), aimed at complex reasoning, agentic workflows, and software development. It's a direct competitor to OpenAI's GPT line, Anthropic's Claude line, and Moonshot AI's Kimi line — not to Midjourney or Sora-style visual generation systems.
MSL's own backstory is worth a paragraph. According to TechCrunch, Silicon Republic, and other outlets, Meta invested roughly $14.3 billion in data-labeling company Scale AI in mid-2025 and brought on its CEO, Alexandr Wang, as Meta's Chief AI Officer to lead this new unit — a move widely reported as stemming from Mark Zuckerberg's dissatisfaction with how far Meta's own Llama line had fallen behind OpenAI and Anthropic. Muse Spark is the first model family MSL has shipped, officially described as a "ground-up overhaul" of Meta's AI stack, with Zuckerberg framing it as the first step toward a "personal superintelligence" vision. There's a strategic pivot worth flagging here too: Llama built its reputation on open weights, but Muse Spark has been closed-weights from version 1.0 onward — no Hugging Face download, no self-hosting, no fine-tuning. That's a clear reversal of Meta's prior open-weight positioning, and we'll come back to the trade-offs of that choice later.
Specifically for 1.2, the model accepts text, images, video, audio, and PDF documents as input and returns text, with a context window of 1,048,576 tokens (roughly 1M). Muse Code, launched the same day, is Meta's first terminal-based coding agent, purpose-built for "long-horizon software engineering tasks" — which is also why nearly everything new in this 1.2 release is concentrated around coding and agentic capability rather than general-purpose chat.
2. Version history: from April's "rebuild from scratch" to the August 5 1.2 release
The timeline matters for understanding how significant this release is. Muse Spark 1.0 debuted on April 8, 2026 — MSL's first product, and the moment Meta's flagship model line pivoted from open weights to closed, commercial API access. Muse Spark 1.1 followed on July 9, 2026, opening paid access through the Meta Model API. According to TechTimes, that launch priced the model at roughly one-quarter of contemporaneous Anthropic and OpenAI rates, while scoring 71 on an independent coding benchmark at about one-third the cost of rival models — the origin of Muse Spark's early reputation for value. Worth noting: when 1.1 first opened API access, it was limited to developers in the United States, not a global rollout.
Muse Spark 1.2 shipped on August 5, 2026, alongside Muse Code (currently in beta) — MSL's third model release in just four months. The gap between 1.0 and 1.1 was roughly three months; the gap between 1.1 and 1.2 was under one month. That accelerating cadence is itself worth recording as a competitive signal: it suggests Meta has made coding agents its current top priority, and is releasing at a noticeably faster clip than the industry average to close ground on Anthropic, OpenAI, and Moonshot AI.
3. Highlight one: cost-efficiency — one of the best "intelligence per dollar" options at its tier
This is Muse Spark 1.2's most defensible strength. Per Artificial Analysis, it scores 54 on the Intelligence Index, up from 51 for 1.1 and 43 for 1.0 — an 11-point climb in four months. Placed against its peers: GPT-5.5 (xhigh) sits at 55 and Grok 4.5 at 54, roughly the same tier. Above that sits the current top cluster — Claude Opus 5 (61), Claude Fable 5 (60), GPT-5.6 Sol (59), and Kimi K3 (57). In other words, Muse Spark 1.2 hasn't cracked the top tier on raw intelligence, but it's now sitting right alongside GPT-5.5 and Grok 4.5, and the gap to the flagship cluster is narrowing.
What makes it worth singling out is cost per unit of intelligence. Artificial Analysis's per-task cost figures, computed against each vendor's own pricing: Muse Spark 1.2 at roughly $0.40, GPT-5.6 Terra (max) at $0.51, Kimi K3 (max) at $0.86, and GPT-5.5 (xhigh) at $1.18. Within that entire intelligence cluster, only Grok 4.5 (high, $0.37) and GPT-5.6 Sol (medium, $0.39) come in cheaper. In other words, if your criterion is "how much intelligence do I get per dollar," Muse Spark 1.2 currently sits near the front of the pack — and this is a measurable claim, not a marketing gloss.
The official API pricing backs this up: $1.25 per million input tokens and $4.25 per million output tokens on the standard tier, with cached input dropping as low as $0.15 per million tokens, alongside the 1M-token context window. That's still meaningfully cheaper than Anthropic's and OpenAI's flagship tiers, continuing the "roughly a quarter of the price" positioning Muse Spark established with 1.1.
4. Highlight two: Muse Code's architecture — persistent background agents plus isolated git-worktree sub-agents
If pricing is the easiest highlight to quantify, Muse Code's architecture is the one most worth an engineer's attention — this isn't a thin CLI wrapper, it makes a handful of genuinely different design choices. According to Meta's own blog and VentureBeat's coverage, Muse Code is a terminal coding agent built for "long-horizon software engineering tasks," able to plan, implement, and validate complex, multi-file changes across large repositories.
Its core design idea is persistent async background agents: Meta describes four background observer agents that stay alive for the entire session, rather than being spun up fresh for each individual task and discarded, as most comparable tools do. That means these background agents can keep working on next steps, avoid re-gathering context that's already been collected, and decide on their own when to report back to the main agent. This is a real architectural difference, not marketing packaging: most coding-agent tools re-scan the codebase and rebuild context every time they spin up a new sub-task on a long-running job. Muse Code's session-persistent design is meant to eliminate a meaningful chunk of that redundant overhead.
For particularly large jobs, Muse Code can fan work out to multiple parallel sub-agents, each running in its own isolated git
worktree, capped at roughly core-count-minus-two. In one reported test, this let the system build six feature modules for a game
project simultaneously with zero collisions. Muse Code also ships with a local event log supporting "replay-exact, restart-safe"
execution, plus /plan, /grill, and /goal commands for steering how the agent plans and executes work.
It's worth placing this design against the rest of the field. Claude Code's differentiation is mostly about surface area — a terminal CLI, VS Code and Cursor extensions, a JetBrains plugin, a desktop app, a web app, Remote Control for steering a local session from a phone, Channels bridging Telegram/Discord/iMessage, and Routines for cloud-hosted recurring tasks. Muse Code has none of that yet. But on the specific question of "how should multiple agents coordinate within a single session," it offers a genuinely different answer from both Claude Code and Codex CLI — that's the real highlight here, not an imitation of an existing terminal tool.
5. Highlight three: the benchmark curve itself — two tiers in three months, closing in on Claude Opus 5
Looking purely at Meta's own reported curve, the improvement is substantial. On Terminal-Bench 2.1, Meta reports 76.2 for 1.1 climbing to 82.9 for 1.2; on DeepSWE v1.1, 53.0 for 1.1 climbing to 59.3 for 1.2. Meta's own launch chart positions Muse Spark 1.2 as second place behind Claude Opus 5 across Terminal-Bench 2.1, DeepSWE v1.1, and Meta's internal coding benchmark. On Terminal-Bench 2.1 specifically, Meta's comparison lists Claude Opus 5 at 86.7, Muse Spark 1.2 at 82.9, GPT-5.6 Terra (running Codex) at 81.8, and Grok Build at 81.6 — if that ordering holds, Muse Spark 1.2 does edge out GPT-5.6 Terra/Codex and Grok Build. For reference, the highest independently verified score any model has posted on Terminal-Bench 2.1 to date is 83.8, from Claude Fable 5 running inside Claude Code.
The DeepSWE v1.1 comparison tells a slightly different story: Muse Spark 1.2 at 59.3, Claude Opus 5 at 65.0, and OpenAI Codex at 64.8 — a roughly 5-6 point gap to the top two, wider than the Terminal-Bench 2.1 gap. Taken together, the fair read is: Muse Spark 1.2 made a genuine two-tier jump in four months and now sits near the leading pack on some coding benchmarks, but the "second only to Claude Opus 5" framing largely comes from Meta's own launch chart. Whether it holds up needs to be calibrated against the independent data in the next section.
6. Highlight four: native multimodality, a 1M-token window, and it's already usable today
Another concrete strength: Muse Spark 1.2's input formats and access channels are both genuinely open rather than stuck at "internal demo" stage. It natively accepts text, images, video, audio, and PDF documents as input, paired with a 1,048,576-token context window — in principle enough to process a long document, a substantial video clip, or a meaningful slice of a large codebase in a single call.
On access: the Meta Model API uses an OpenAI-SDK-compatible calling convention (base URL api.meta.ai/v1) and documents support for Chat Completions, Responses, and Messages protocol formats — meaning developers already wired up for OpenAI-format or Anthropic Messages-format models can, in principle, plug Muse Spark into an existing toolchain without writing a separate protocol-translation layer. That's the same logic DeepSeek used when it added native Responses API support to work smoothly with Codex. Beyond the Meta Model API, both Muse Spark 1.1 and 1.2 are live on OpenRouter, and according to OpenRouter, the 1.2 launch came with "expanded global access" — a step up from 1.1's initial US-only availability. That's a real opening move, though the specific list of newly covered countries and regions hasn't been published anywhere we could find; we'll come back to that uncertainty in the relay-access section.
7. The cold water that has to be thrown: how big is the self-reported vs. verified gap
None of Meta's self-reported scores for 1.2 have been independently verified yet
The 82.9 (Terminal-Bench 2.1) and 59.3 (DeepSWE v1.1) figures cited above are, as of this writing, numbers from Meta's own launch blog only — neither appears yet on an official verified leaderboard. And the historical record from version 1.1 shows this gap can be substantial.
The 1.1 case is the one we can actually cross-check. Meta's self-reported Terminal-Bench 2.1 score at 1.1's launch was 80.0. The official Terminal-Bench verified leaderboard reproduction came in at 76.2 (±1.2) — even taking the upper bound of 77.4, that's still 2.6 points below Meta's number. Independent evaluator Vals AI ran its own test harness and measured 69.29 — a gap of more than 10 points from Meta's self-reported 80.0. Those three numbers (80.0, 76.2, 69.29) come from three different measurement methodologies, and that alone is a useful reminder: in the coding-agent category, "the score" is highly dependent on the specific test harness and evaluation protocol in use — the same model can post materially different numbers depending on who's running the test.
For 1.2, the pattern repeats: Meta reports 82.9, but as of this writing it hasn't appeared on an official verified leaderboard. According to Vals AI's independent testing, Muse Spark 1.2 ranks just 14th on the common-harness Terminal-Bench leaderboard, but 5th overall on Vals's composite Index — which blends intelligence and cost efficiency — at a remarkably low $0.69 per test. Put together, this suggests Muse Spark 1.2's raw coding accuracy, measured by an independent test harness, probably doesn't crack the top tier; but because it's so cheap, once you fold cost into the ranking alongside accuracy, its overall standing rises substantially. That's exactly why this review keeps emphasizing "benchmark value" rather than "raw benchmark score" — it's the most honest framing the currently available independent data supports.
8. Other real limitations: closed weights, an immature Muse Code, and the privacy cost of the cheap tier
Beyond the benchmark-methodology gap, there are several concrete limitations worth including — a review that only covers highlights wouldn't be a complete one:
- Fully closed weights, no self-hosting: continuing the pivot that started with 1.0, there's no Hugging Face download, no self-hosting, and no fine-tuning — a clear reversal from Llama's open-weight-flagship positioning. Compared to the open-weight models this site has reviewed — DeepSeek V4, Kimi K3, GLM-5.2 — any enterprise scenario that requires private, on-prem deployment simply cannot choose Muse Spark right now.
- Architecture and parameter count are entirely undisclosed: Meta hasn't published whether Muse Spark uses a mixture-of-experts architecture, its parameter count, or a knowledge cutoff date — it's a fully opaque black box, in contrast to vendors like DeepSeek and Moonshot AI that publish technical papers with architecture details.
- Muse Code is still beta, and missing key platform/ecosystem support: at launch it only runs on macOS and Linux, with no Windows build; there's also no VS Code, JetBrains, or Neovim plugin — it's a pure terminal tool without Claude Code's multi-surface ecosystem (IDE plugins, desktop app, web app, remote control, messaging-platform bridges).
- The cheaper "contributor" tier comes with a privacy trade-off: the Meta Model API offers a cheaper muse-spark-1.2-contributor tier ($0.10 input / $0.20 output per million tokens), but in exchange, you agree to let Meta use those prompts and completions to train future models — and that tier also caps you at 60 requests per minute. Multiple independent reviews flag this tier as unsuitable for most professional or enterprise codebases; anything involving private code or business logic you don't want used as training data should stick to the standard tier.
9. Side by side: Muse Spark 1.2 against the current crop of leading models
The two tables below summarize the Intelligence Index / cost-efficiency comparison, and the Terminal-Bench 2.1 coding benchmark comparison specifically. Figures come from Artificial Analysis and vendor disclosures; some Terminal-Bench numbers come from Meta's own launch chart (flagged as such — not independently verified):
| Model | AA Intelligence Index | Cost per task (top tier) |
|---|---|---|
| Claude Opus 5 | 61 | No reliable figure found |
| Claude Fable 5 | 60 | No reliable figure found |
| GPT-5.6 Sol (medium) | 59 | $0.39 |
| Kimi K3 (max) | 57 | $0.86 |
| GPT-5.5 (xhigh) | 55 | $1.18 |
| Muse Spark 1.2 | 54 | $0.40 |
| Grok 4.5 (high) | 54 | $0.37 |
| Model | Terminal-Bench 2.1 | Source / verification status |
|---|---|---|
| Claude Fable 5 (Claude Code) | 83.8% | Highest known independently verified score |
| Claude Opus 5 | 86.7% | Cited from Meta's launch chart, verification status not stated |
| Muse Spark 1.2 | 82.9% | Meta self-reported, not yet on an official verified leaderboard |
| GPT-5.6 Terra (Codex) | 81.8% | Cited from Meta's launch chart |
| Grok Build | 81.6% | Cited from Meta's launch chart |
| Muse Spark 1.1 (independently reproduced) | 76.2% (±1.2) | Official Terminal-Bench verified leaderboard |
| Muse Spark 1.1 (Vals AI independent test) | 69.29% | Vals AI's own test harness |
The second table deliberately lays out all three of Muse Spark 1.1's different scores (Meta's self-reported 80.0, noted in the section above but not repeated in the table; the independently verified 76.2; and Vals AI's own 69.29) as a reminder: many of the other models' scores sitting above or below Muse Spark 1.2 in this table also come from each vendor's own launch materials, and haven't necessarily been independently verified under a consistent methodology either. Treat any single vendor's chart with a healthy amount of caution.
10. What this means for mainland China AI API relay readers
Unlike OpenAI's Astra teaser, which this site covered earlier, Muse Spark 1.2 isn't a model that "hasn't launched yet" — it's already available through paid access on the Meta Model API and is already live on OpenRouter, which means the basic technical conditions for relay-provider integration already exist. A few specifics are worth stating plainly rather than reducing this to a simple yes/no.
First, OpenRouter describes the 1.2 launch as coming with "expanded global access," a step up from 1.1's initial US-only rollout — a real move toward openness. But we found no official source detailing exactly which countries and regions are now covered, so we cannot confirm whether mainland China is explicitly included; any claim that "official access to China has been confirmed" currently lacks a verifiable source. Second, Meta's consumer-facing AI products — the Meta AI app, WhatsApp, and Instagram's AI features — have long not operated in mainland China, which is a reasonable but still speculative data point: even if the Meta Model API is nominally open "globally," whether it's actually reachable and stable from mainland China remains uncertain, and we won't draw an unverified conclusion on the reader's behalf here.
Third, from a technical standpoint, because Muse Spark 1.2 is already on OpenRouter — itself an upstream aggregation channel many
relay providers already draw from — any relay provider already offering OpenRouter models is, in principle, technically positioned to
add the meta/muse-spark-1.2 model ID to its own catalog without a separate direct business relationship with Meta. That
said, as of this writing we found no relay provider that has publicly announced support for Muse Spark 1.2 — if you see marketing
claiming otherwise, verify the exact model ID, actual regional access, and real pricing directly with that vendor rather than taking a
landing page at face value. Fourth, because the Meta Model API is compatible with the OpenAI SDK calling convention and documents
support for Chat Completions, Responses, and Messages protocol formats, any relay provider that does add it should, in principle, be
able to reuse existing gateway configuration built for OpenAI-format or Claude Messages-format models (via tools like LiteLLM or Claude
Code Router) without building a separate protocol-translation layer — a real practical upside for developers already running a
multi-model gateway setup.
11. Who should use it now, and who should wait
- Developers highly sensitive to token cost, working on coding/agentic workflows → Muse Spark 1.2's cost-per-unit- of-intelligence is currently near the front of the pack at its tier — worth including directly in your evaluation.
- Engineers curious about new coding-agent architectures → Muse Code's persistent background-agent plus isolated git-worktree sub-agent design is a genuine architectural difference; worth trying hands-on against Claude Code and Codex CLI.
- Enterprise scenarios requiring private, on-prem deployment → Muse Spark is fully closed-weights with no self-hosting or fine-tuning, and currently doesn't fit here at all; consider open-weight alternatives like DeepSeek, Kimi, or GLM.
- Developers who need Windows support or deep IDE integration (VS Code/JetBrains plugins) → Muse Code doesn't cover these yet at launch; worth waiting a few more versions, or sticking with a tool like Claude Code that already has a mature multi-surface ecosystem.
- Readers inclined to take a vendor's own launch chart at face value → Cross-reference the independent verification data in section 7, especially the gap between Meta's self-reported Terminal-Bench 2.1 score and the Vals AI / official leaderboard numbers, before treating "second only to Claude Opus 5" as settled fact.
- Mainland China teams accustomed to relay-provider access to overseas models → There's currently no solid evidence that mainland China is explicitly included in official "global" access, and no relay provider has publicly announced support — worth watching, but verify directly with a specific vendor before committing.
12. Conclusion
Meta Muse Spark 1.2 is a clearly positioned multimodal reasoning and coding model — not the image or video generation tool its name might suggest to some readers. In this August 5, 2026 release, its genuinely defensible highlights cluster around four things: first, cost-efficiency — by Artificial Analysis's math, it's one of the cheapest options per task among models at its intelligence tier; second, the architectural innovation in Muse Code — persistent background observer agents paired with isolated git-worktree parallel sub-agents, a genuinely different technical approach from both Claude Code and Codex CLI, not an imitation; third, the pace of improvement itself — climbing from 43 to 54 on the Intelligence Index in four months, closing in on Claude Opus 5 on some coding benchmarks; and fourth, native multimodal input and a 1M-token context window, already accessible through a relatively open combination of the Meta Model API and OpenRouter.
But it's equally important to state plainly: Meta's self-reported benchmark numbers — especially the 82.9 on Terminal-Bench 2.1 — remain unverified as of this writing, and the historical record from 1.1 shows the gap between Meta's self-reported scores and third-party reproductions can exceed 10 points. Muse Spark is fully closed-weights and can't be self-hosted, a clear reversal from Meta's earlier Llama-era openness. Muse Code is still in beta, missing Windows support and an IDE plugin ecosystem. And the cheaper "contributor" pricing tier requires giving up rights to your data for model training, making it unsuitable for professional use. For mainland China AI API relay readers, the most honest conclusion is: the model itself is technically ready to be integrated (it's live on OpenRouter and supports standard protocol formats), but whether mainland China is explicitly covered by official global access, and whether any relay provider has formally added support, both remain unverifiable as of this writing — worth watching, but verify with a specific vendor before acting, rather than trusting anyone's marketing copy.
Put together, this means
If you care about cost-efficiency and new architectural ideas for coding/agentic workflows, Muse Spark 1.2 and Muse Code are worth trying hands-on right now. If you need private on-prem deployment, deep Windows/IDE integration, or you take vendor-reported benchmarks at face value, this isn't the right moment yet.
- Cost-sensitive developers working on coding/agentic tasks → Worth comparing hands-on now; the value data currently holds up.
- Enterprise on-prem deployment needs → Not currently supported; look at open-weight alternatives instead.
- Mainland China relay users → The model is technically ready for integration, but official access scope and relay-provider support both still need direct verification — don't trust unverified "already supported" claims.