
Meta Muse Spark: Impressive at Health, Weak at Coding – and a Strategic Problem
TL;DR: „Muse Spark is Meta's best model ever and it's free. It leads on health benchmarks and scientific reasoning, but falls dramatically behind on coding (59 vs. 75 GPT-5.4) and agentic tasks. The real elephant in the room: the model is closed-source – a break with Meta's open-source DNA."
— Till FreitagThe Key Takeaway in 30 Seconds
On April 8, 2026, Meta unveiled Muse Spark – the first model from the new Meta Superintelligence Labs (MSL) under the leadership of Alexandr Wang. Nine months of rebuilding, new stack, new architecture.
The headlines sound impressive: top 5 in the Artificial Analysis Intelligence Index, best model for medical reasoning, free for all 3+ billion Meta users.
But a closer look reveals a more nuanced picture.
What Muse Spark Can Actually Do
Health & Medical AI: In a League of Its Own
This is where Meta delivered. 42.8 on HealthBench Hard – better than GPT-5.4 (40.1) and twice as good as Gemini 3.1 Pro (20.6). Over 1,000 physicians contributed to the training data.
This isn't a coincidence – it's strategy: Meta has 3 billion users on WhatsApp, Instagram, and Facebook. An AI assistant that can reliably answer health questions is a massive retention lever.
Multimodal Vision: Strong, But Not Number One
80.5% on MMMU-Pro (Gemini scores 82.4%). On CharXiv Reasoning – chart and data comprehension – Muse Spark leads with 86.4 over GPT-5.4 (82.8). Anyone working heavily with visual data will find a strong tool here.
Scientific Reasoning: The Contemplating Mode
Muse Spark's killer feature is the Contemplating mode: instead of pushing a single model to think harder, it orchestrates multiple agents in parallel. The result: 50.2% on Humanity's Last Exam – better than GPT-5.4 Pro (43.9%) and Gemini Deep Think (48.4%).
Where Muse Spark Falls Short
Coding: Not Even Close to Competitive
This is where it gets uncomfortable. Terminal-Bench 2.0: 59.0 – while GPT-5.4 scores 75.1 and Gemini hits 68.5. This isn't a small gap – it's a different league.
For anyone using AI for programming – and that's an ever-growing number of developers and "vibe coders" – Muse Spark is simply not an option. Claude and GPT remain unchallenged here.
Agentic Tasks: Not Ready for Autonomous Work
GDPval-AA: 1,444 ELO vs. GPT-5.4 (1,674) and Claude Opus 4.6 (1,607). When an AI model needs to independently execute multi-step workflows – filling spreadsheets, navigating websites, managing documents – Muse Spark isn't reliable enough.
Abstract Reasoning: The Biggest Blind Spot
ARC-AGI-2: 42.5 vs. GPT-5.4 (76.1) and Gemini (76.5). This isn't close – it's less than half. On novel pattern recognition tasks that require genuine generalization, Muse Spark breaks down.
The Benchmark Table
| Benchmark | Muse Spark | GPT-5.4 | Claude Opus 4.6 | Gemini 3.1 Pro |
|---|---|---|---|---|
| AI Analysis Index | 52 | 57 | 53 | 57 |
| Humanity's Last Exam | 50.2% | 43.9% | — | 48.4% |
| HealthBench Hard | 42.8 | 40.1 | — | 20.6 |
| CharXiv Reasoning | 86.4 | 82.8 | — | 80.2 |
| MMMU-Pro (Vision) | 80.5% | — | — | 82.4% |
| Terminal-Bench (Coding) | 59.0 | 75.1 | — | 68.5 |
| ARC-AGI-2 | 42.5 | 76.1 | — | 76.5 |
| GDPval-AA (Agentic) | 1,444 | 1,674 | 1,607 | — |
| Price | Free | Subscription | Subscription | Freemium |
The Elephant in the Room: Closed Source
And here's where it gets strategically interesting. Meta waved the open-source flag for years. Llama was the declared countermodel to OpenAI's and Google's closed-source approach. Zuckerberg positioned open source as a moral imperative.
Muse Spark is closed-source.
Yes, Meta has announced it will release open-source weights "in the future." But there's no timeline. And the fact that the company's best model sits behind a closed API sends a clear signal: When it comes to true frontier performance, Meta prioritizes control over openness.
This isn't a criticism of the decision per se – OpenAI and Anthropic do the same. But it undermines the narrative Meta used to differentiate itself for years.
What This Means for the AI Race
1. Meta Is Playing a Different Game
While OpenAI builds the consumer super-app and Anthropic builds the developer OS, Meta focuses on distribution. 3+ billion users who can use Muse Spark for free via WhatsApp, Instagram, and Facebook – that's a moat no startup can replicate.
2. The Health Focus Is Clever
Meta isn't positioning AI as a productivity tool but as a personal health advisor. This sounds like a nice feature, but it's potentially a paradigm shift: if users start asking their health questions to an AI assistant on WhatsApp instead of Google, a multi-billion-dollar market shifts.
3. Coding Remains Meta's Achilles' Heel
Here's the dilemma: the developer community – the multipliers who build ecosystems – needs coding capabilities. And that's exactly where Muse Spark is weakest. Until Meta closes this gap, Claude/GPT remains the developer tool of choice.
Our Assessment
Muse Spark is an impressive first model from Meta Superintelligence Labs. The health focus is strategically smart, the Contemplating mode is technically innovative, and the free availability puts pressure on the competition.
But for professional use cases – coding, agent workflows, abstract problem-solving – Muse Spark is not competitive as of today. It's a consumer model with frontier ambitions, not a frontier model with consumer reach.
The biggest contradiction remains the closed-source decision. Meta has to choose: does it want to be the company that democratizes AI – or the company that locks 3 billion users into a closed AI ecosystem? Both at the same time doesn't work long-term.
We'll integrate Muse Spark into our tool comparisons in the coming weeks and monitor its development. Meta has the resources to close the coding gap. The question is whether they want to – or whether they'd rather become the world's best health AI provider.
Sources: Meta AI Blog, Artificial Analysis Intelligence Index v4.0, FelloAI Benchmark Analysis








