Meta built the most expensive AI lab on paper. Then it released Muse Spark -- a model that sits comfortably in the top five and nowhere near the top one.
On paper, Muse Spark is not bad. [Artificial Analysis](https://artificialanalysis.ai) ranks it fifth overall, behind only Gemini 3.1 Pro, GPT-5.4, Claude Opus 4.6, and Claude Mythos Preview. It posts 89.5% on [GPQA Diamond](https://benchmark.space/benchmark/gpqa-diamond), 77.4% on [SWE-bench Verified](https://benchmark.space/benchmark/swe-bench-verified), and a respectable 80.5% on MMMU Pro. It even claims SOTA on niche benchmarks like CharXiv (86.4%) and HealthBench Hard (42.8%).
But peel back the top-line numbers and the gaps open up.
| Benchmark | Muse Spark | Leader | Gap |
| GPQA Diamond | 89.5% | 94.6% (Claude Mythos) | 5.1 pts |
| SWE-bench Verified | 77.4% | 93.9% (Claude Mythos) | 16.5 pts |
| ARC-AGI 2 | 42.5% | 85.0% (Claude Opus 4.6) | 42.5 pts |
| FrontierMath Tier 4 | 15.0% | 38.0% (GPT-5.4 Pro) | 23.0 pts |
A 16-point gap on [SWE-bench Verified](https://benchmark.space/benchmark/swe-bench-verified) is not a rounding error. A 42-point gap on [ARC-AGI 2](https://benchmark.space/benchmark/arc-agi-2) is a different tier of model.
The Agentic Hole
The sharper problem is agentic performance. Muse Spark does not appear on the leaderboards for OSWorld, WebArena, Mind2Web, or BrowseComp -- the benchmarks that measure whether a model can actually do things, not just answer questions.
As Sam Witteveen noted in his breakdown of the release: "Unfortunately for Meta it does not look like the model stands out very strong on agentic performance." — [Sam Witteveen, 9:17](https://www.youtube.com/watch?v=7vkybiVRSm0&t=557)
This is the same pattern that separates the top three labs from everyone else. GPT-5.4 dominates [Mind2Web](https://benchmark.space/benchmark/mind2web) at 92.8%. Claude Mythos Preview leads [OSWorld](https://benchmark.space/benchmark/osworld) at 79.6%. These are agent benchmarks -- the ones that correlate with actual product capability. Muse Spark is absent from the conversation entirely.
The Alexander Wang Irony
One benchmark result crystallizes the awkwardness. On [Humanity's Last Exam](https://benchmark.space/benchmark/humanitys-last-exam) (with tools), Muse Spark sits at 58.0% -- tied for last place among frontier models. Humanity's Last Exam was created by Scale AI. Scale AI's founder, Alexander Wang, now runs Meta Superintelligence Labs.
"Not only is their model winning on only three of the benchmarks, they are also actually last on three of the benchmarks including Humanity's Last Exam here with tools, which is the benchmark that Alexander Wang's own company Scale AI actually created." — [Sam Witteveen, 7:15](https://www.youtube.com/watch?v=7vkybiVRSm0&t=435)
Being last on a benchmark your own lab director's company designed is the kind of detail that writes itself.
The Proprietary Pivot
Perhaps the most consequential shift is not the benchmark scores but the licensing. Muse Spark is closed. No open weights. No API. This is a hard break from the Llama era, where Meta positioned itself as the open-weight champion and Llama models became the default starting point for the community.
That community has already moved on. [Qwen](https://benchmark.space/model/qwen-3.5-9b) and Gemma 4 now dominate the open-weight landscape that Llama once owned. Qwen 3.5 9B scores 92.5% on [AIME](https://benchmark.space/benchmark/aime) -- an open model a fraction of Muse Spark's size matching frontier-level math. Meta walked away from the exact segment where it still had leverage.
The Acquisition Strategy
Meta's spend is staggering. The Manas acquisition reportedly ran $1-2 billion for a multi-user agent system. Daniel Gross and Nat Friedman's startup was another large acquisition. Scale AI's Wang came with his own deal. The talent roster is deep. But the output so far is a model that Witteveen summarizes as sitting "ahead of Claude Sonnet, ahead of the GLM 5.1, MiniMax 2.7. And really only behind the models from the top three proprietary labs." — [Sam Witteveen, 8:16](https://www.youtube.com/watch?v=7vkybiVRSm0&t=496)
"Only behind the top three" sounds like fourth place. In a winner-take-most market, fourth place is a polite way of saying not competitive.
What Money Cannot Buy
The lesson is not that money is irrelevant. Money buys compute, talent, and time. But the labs that lead -- Anthropic, OpenAI, Google -- have something money struggles to replicate: years of compounding organizational knowledge about training runs, data pipelines, and the fragile alchemy of turning a research insight into a production model. You cannot acquire your way into that. You have to build it.
Muse Spark is a competent debut from a reassembled lab. It is not a threat to the frontier. And until Meta can show agentic capability and close the reasoning gaps, the billions spent will look more like an expensive admission ticket than a winning hand.