Meta Muse Spark: Impressive at Health, Weak at Coding – and a Strategic Problem

    Meta Muse Spark: Impressive at Health, Weak at Coding – and a Strategic Problem

    13. April 20264 min read
    Till Freitag

    TL;DR:Muse Spark is Meta's best model ever and it's free. It leads on health benchmarks and scientific reasoning, but falls dramatically behind on coding (59 vs. 75 GPT-5.4) and agentic tasks. The real elephant in the room: the model is closed-source – a break with Meta's open-source DNA."

    Till Freitag

    The Key Takeaway in 30 Seconds

    On April 8, 2026, Meta unveiled Muse Spark – the first model from the new Meta Superintelligence Labs (MSL) under the leadership of Alexandr Wang. Nine months of rebuilding, new stack, new architecture.

    The headlines sound impressive: top 5 in the Artificial Analysis Intelligence Index, best model for medical reasoning, free for all 3+ billion Meta users.

    But a closer look reveals a more nuanced picture.


    What Muse Spark Can Actually Do

    Health & Medical AI: In a League of Its Own

    This is where Meta delivered. 42.8 on HealthBench Hard – better than GPT-5.4 (40.1) and twice as good as Gemini 3.1 Pro (20.6). Over 1,000 physicians contributed to the training data.

    This isn't a coincidence – it's strategy: Meta has 3 billion users on WhatsApp, Instagram, and Facebook. An AI assistant that can reliably answer health questions is a massive retention lever.

    Multimodal Vision: Strong, But Not Number One

    80.5% on MMMU-Pro (Gemini scores 82.4%). On CharXiv Reasoning – chart and data comprehension – Muse Spark leads with 86.4 over GPT-5.4 (82.8). Anyone working heavily with visual data will find a strong tool here.

    Scientific Reasoning: The Contemplating Mode

    Muse Spark's killer feature is the Contemplating mode: instead of pushing a single model to think harder, it orchestrates multiple agents in parallel. The result: 50.2% on Humanity's Last Exam – better than GPT-5.4 Pro (43.9%) and Gemini Deep Think (48.4%).


    Where Muse Spark Falls Short

    Coding: Not Even Close to Competitive

    This is where it gets uncomfortable. Terminal-Bench 2.0: 59.0 – while GPT-5.4 scores 75.1 and Gemini hits 68.5. This isn't a small gap – it's a different league.

    For anyone using AI for programming – and that's an ever-growing number of developers and "vibe coders" – Muse Spark is simply not an option. Claude and GPT remain unchallenged here.

    Agentic Tasks: Not Ready for Autonomous Work

    GDPval-AA: 1,444 ELO vs. GPT-5.4 (1,674) and Claude Opus 4.6 (1,607). When an AI model needs to independently execute multi-step workflows – filling spreadsheets, navigating websites, managing documents – Muse Spark isn't reliable enough.

    Abstract Reasoning: The Biggest Blind Spot

    ARC-AGI-2: 42.5 vs. GPT-5.4 (76.1) and Gemini (76.5). This isn't close – it's less than half. On novel pattern recognition tasks that require genuine generalization, Muse Spark breaks down.


    The Benchmark Table

    BenchmarkMuse SparkGPT-5.4Claude Opus 4.6Gemini 3.1 Pro
    AI Analysis Index52575357
    Humanity's Last Exam50.2%43.9%48.4%
    HealthBench Hard42.840.120.6
    CharXiv Reasoning86.482.880.2
    MMMU-Pro (Vision)80.5%82.4%
    Terminal-Bench (Coding)59.075.168.5
    ARC-AGI-242.576.176.5
    GDPval-AA (Agentic)1,4441,6741,607
    PriceFreeSubscriptionSubscriptionFreemium

    The Elephant in the Room: Closed Source

    And here's where it gets strategically interesting. Meta waved the open-source flag for years. Llama was the declared countermodel to OpenAI's and Google's closed-source approach. Zuckerberg positioned open source as a moral imperative.

    Muse Spark is closed-source.

    Yes, Meta has announced it will release open-source weights "in the future." But there's no timeline. And the fact that the company's best model sits behind a closed API sends a clear signal: When it comes to true frontier performance, Meta prioritizes control over openness.

    This isn't a criticism of the decision per se – OpenAI and Anthropic do the same. But it undermines the narrative Meta used to differentiate itself for years.


    What This Means for the AI Race

    1. Meta Is Playing a Different Game

    While OpenAI builds the consumer super-app and Anthropic builds the developer OS, Meta focuses on distribution. 3+ billion users who can use Muse Spark for free via WhatsApp, Instagram, and Facebook – that's a moat no startup can replicate.

    2. The Health Focus Is Clever

    Meta isn't positioning AI as a productivity tool but as a personal health advisor. This sounds like a nice feature, but it's potentially a paradigm shift: if users start asking their health questions to an AI assistant on WhatsApp instead of Google, a multi-billion-dollar market shifts.

    3. Coding Remains Meta's Achilles' Heel

    Here's the dilemma: the developer community – the multipliers who build ecosystems – needs coding capabilities. And that's exactly where Muse Spark is weakest. Until Meta closes this gap, Claude/GPT remains the developer tool of choice.


    Our Assessment

    Muse Spark is an impressive first model from Meta Superintelligence Labs. The health focus is strategically smart, the Contemplating mode is technically innovative, and the free availability puts pressure on the competition.

    But for professional use cases – coding, agent workflows, abstract problem-solving – Muse Spark is not competitive as of today. It's a consumer model with frontier ambitions, not a frontier model with consumer reach.

    The biggest contradiction remains the closed-source decision. Meta has to choose: does it want to be the company that democratizes AI – or the company that locks 3 billion users into a closed AI ecosystem? Both at the same time doesn't work long-term.

    We'll integrate Muse Spark into our tool comparisons in the coming weeks and monitor its development. Meta has the resources to close the coding gap. The question is whether they want to – or whether they'd rather become the world's best health AI provider.


    Sources: Meta AI Blog, Artificial Analysis Intelligence Index v4.0, FelloAI Benchmark Analysis

    TeilenLinkedInWhatsAppE-Mail

    Related Articles

    Abstract visualization of AI transformation: chaotic data structures channeled through glowing neural pathways into ordered architecture
    May 10, 20265 min

    AI Transformation: Roadmap, Change Management & Implementation Phases for Companies

    AI Transformation is more than handing out ChatGPT licenses. An honest roadmap with five phases, change management princ

    Read more
    Geopolitical AI landscape between western and eastern technologyDeep Dive
    April 13, 20268 min

    China's AI Offensive: From Hunter Alpha to DeepSeek V4 on Huawei Chips

    An anonymous 1T model, a DeepSeek mix-up, and the reveal that Xiaomi was behind it. Meanwhile, DeepSeek V4 on Huawei chi

    Read more
    The AI Race in 43 Milestones: The Complete OpenAI vs. Anthropic Timeline
    April 11, 20264 min

    The AI Race in 43 Milestones: The Complete OpenAI vs. Anthropic Timeline

    From GPT-4o to Project Glasswing: Every acquisition, model launch, and product release from OpenAI and Anthropic on an i

    Read more
    Google's $185 Billion Bet: How Gemini 3.1 Pro, Vertex AI, and the Largest Infrastructure Offensive in Tech History Are Reshaping the AI Race
    April 11, 20267 min

    Google's $185 Billion Bet: How Gemini 3.1 Pro, Vertex AI, and the Largest Infrastructure Offensive in Tech History Are Reshaping the AI Race

    Alphabet is investing up to $185 billion in AI infrastructure – more than the GDP of 140 countries. Gemini 3.1 Pro doubl

    Read more
    OpenAI Buys a TV Show. Anthropic Builds the Future of Software. And Google? It's Playing a Different Game Entirely.
    April 11, 20266 min

    OpenAI Buys a TV Show. Anthropic Builds the Future of Software. And Google? It's Playing a Different Game Entirely.

    OpenAI buys TBPN, a Jony Ive hardware startup, and builds a desktop superapp. Anthropic turns Claude into a Developer OS

    Read more
    Stylized number 5 made of orange ribbons and gears – cover for Claude Sonnet 5
    June 30, 20263 min

    Claude Sonnet 5: Agentic AI Goes Mainstream

    Anthropic ships Claude Sonnet 5 – a Sonnet model that gets close to Opus 4.8 performance at a fraction of the price. Thi

    Read more
    AI Benchmarks Explained: Arena, SWE-Bench, AutomationBench & Co.
    June 26, 20266 min

    AI Benchmarks Explained: Arena, SWE-Bench, AutomationBench & Co.

    How do AI benchmarks actually work – from LMArena to SWE-Bench to Zapier's AutomationBench? A tour of Elo rankings, stat

    Read more
    GLM-5.2 vs. Kimi K2.7 Code – split-screen illustration with Z-letter mark and crescent moon symbol
    June 21, 20267 min

    GLM-5.2 vs. Kimi K2.7 Code: Two Open-Weight Releases in One Week – Two Very Different Bets

    Within four days in June 2026, Z.ai (GLM-5.2) and Moonshot AI (Kimi K2.7 Code) shipped their next-generation open-weight

    Read more
    Gemma 4 12B Coder running locally on a developer laptop – code symbols streaming from a 12B chip
    June 15, 20264 min

    Gemma 4 12B Coder: Local Code Generation Becomes the Default

    Google ships the Gemma 4 12B Coder — the specialized coding variant of the Gemma 4 stack. 12B parameters in GGUF format,

    Read more