Meta Muse Spark launched in April 2026 scoring 52 on the Artificial Analysis Intelligence Index v4.0, and it was the model that put Meta back in the frontier conversation. Two upgrades later the line has moved a long way; Muse Spark 1.2 now scores 57 on index v4.1.1 and ships with a terminal coding agent. Built by Meta Superintelligence Labs (MSL), the team led by Alexandr Wang, Muse Spark was the result of a nine-month ground-up rebuild of Meta’s entire AI infrastructure, and it is still free to use.
But raw rankings only tell part of the story. Muse Spark beat every competitor on medical and health benchmarks at launch, ranked second in multimodal vision understanding, and introduced a unique Contemplating mode that runs multiple AI agents in parallel. Coding and agentic work were its clear weak spots in April, and those are exactly the gaps Meta has spent the months since closing. This guide breaks down what Muse Spark does well and where it fell short. It also covers how far the model has come since, how it compares to the other leading models, and how to start using it.
Update, August 7, 2026. This page documents the original April 2026 Muse Spark release, and the benchmark table below is that launch comparison. Meta has shipped twice since, Muse Spark 1.1 on July 9, 2026 with its first paid API, and Muse Spark 1.2 with the Muse Code terminal agent on August 5, 2026. The rivals have all moved too. Claude Opus 5 replaced Opus 4.6 and leads the Intelligence Index at 63, and OpenAI’s frontier family is now GPT-5.6 in Sol, Terra and Luna tiers, not GPT-5.4 or GPT-5.5. Where a section below reports launch-day figures, it says so.
The Key Takeaways
- The April 2026 release scored 52 on Artificial Analysis Intelligence Index v4.0. On the current v4.1.1 board, Muse Spark 1.2 scores 57 and the original 1.0 is no longer listed
- Best-in-class health AI, scoring 42.8 on HealthBench Hard, beating GPT-5.4 (40.1) and Gemini 3.1 Pro (20.6)
- Three reasoning modes including Contemplating, which scored 50.2% on Humanity’s Last Exam, beating both GPT-5.4 Pro (43.9%) and Gemini Deep Think (48.4%)
- Completely free at meta.ai and in the Meta AI app. Developer access is no longer a private preview; the Meta Model API has been public and paid since July 9, 2026
- Coding and agentic work were the launch weak spots (Terminal-Bench 59.0, GDPval-AA 1,444 ELO). Meta has closed most of that gap since; on its own charts Muse Spark 1.2 scores 82.9% on Terminal-Bench 2.1, second to Claude Opus 5’s 86.7%
What Is Meta Muse Spark?
Meta Muse Spark is a natively multimodal reasoning model developed by Meta Superintelligence Labs, Meta’s elite AI research division. Internally codenamed “Avocado,” the model was built over nine months after Meta scrapped its previous approach and rebuilt the entire AI stack from scratch, including new infrastructure, architecture, and data pipelines.
Unlike Llama 4, which is open-source, Muse Spark is currently a closed model. Meta has stated it plans to release open-source weights in the future, but no timeline has been announced. The model powers Meta AI across Meta’s platforms, and rolled out to Facebook, Instagram and WhatsApp over the weeks after launch. If you would rather not have the assistant in your social apps at all, see how to turn off Meta AI on every Meta surface.
Muse Spark accepts text, image, and voice inputs but currently produces text-only output. It supports tool-use, visual chain of thought reasoning, and multi-agent orchestration. According to Meta, the model was trained using over 10x less compute than Llama 4 Maverick while achieving significantly better performance. For Meta’s image side, see our guide to the Meta AI image generator.
Meta Muse Spark Benchmarks: How It Compares
The benchmark data tells a nuanced story. Muse Spark was genuinely competitive with frontier models in several areas while having clear gaps in others. The table below is the April 2026 launch comparison, run against the models that were current then, GPT-5.4, Claude Opus 4.6 and Gemini 3.1 Pro. We have left it as the dated record it is rather than restaging it, because it is what Meta and Artificial Analysis published at launch. For today’s board, see the update note above.
| Benchmark | Muse Spark | GPT-5.4 | Claude Opus 4.6 | Gemini 3.1 Pro |
|---|---|---|---|---|
| Artificial Analysis Index | 52 | 57 | 53 | 57 |
| Humanity’s Last Exam | 50.2% (Contemplating) | 43.9% Pro | — | 48.4% Deep Think |
| HealthBench Hard | 42.8 | 40.1 | — | 20.6 |
| CharXiv Reasoning | 86.4 | 82.8 | — | 80.2 |
| MMMU-Pro (Vision) | 80.5% | — | — | 82.4% |
| DeepSearchQA | 74.8 | — | — | 69.7 |
| ARC-AGI-2 | 42.5 | 76.1 | — | 76.5 |
| Terminal-Bench (Coding) | 59.0 | 75.1 | — | 68.5 |
| GDPval-AA (Agentic) | 1,444 ELO | 1,674 | 1,607 | — |
| FrontierScience Research | 38.3% | 36.7% | — | 23.3% |
| IPhO 2025 Theory | 82.6 | 93.5 | — | 87.7 |
| ZeroBench (Visual) | 33.0 | 41.0 | — | 29.0 |
| MedXpertQA | 78.4 | 77.1 | — | 81.3 |
| Price | Free | Subscription | Subscription | Free tier + paid |
Sources: Meta AI blog, Artificial Analysis Intelligence Index v4.0. Figures as published at the April 2026 launch; the index is now on v4.1.1.
One standout metric was token efficiency. Muse Spark used just 58 million output tokens to complete the full Intelligence Index evaluation at launch. That was comparable to Gemini 3.1 Pro but far less than Claude Opus 4.6 (157M) and GPT-5.4 (120M). Efficient token use typically translates to faster responses and lower computational costs.
Where Meta Muse Spark Excels
Muse Spark is not trying to be the best at everything. But in a few specific domains, it either leads or is very close to the top.
Health and Medical AI
This is Muse Spark’s strongest area. With a 42.8 score on HealthBench Hard, it outperforms every other model tested, including GPT-5.4 (40.1) and Gemini 3.1 Pro (20.6). Meta collaborated with over 1,000 physicians to curate specialized training data for health-related queries. The model can display interactive nutritional breakdowns, explain exercise biomechanics, and provide factual health information with visual aids.
Multimodal Vision Understanding
Muse Spark scores 80.5% on MMMU-Pro, making it the second-most capable multimodal model behind Gemini 3.1 Pro (82.4%). On CharXiv Reasoning, which tests figure and chart understanding, Muse Spark leads with 86.4, ahead of GPT-5.4 (82.8) and Gemini (80.2). If your work involves analyzing images, charts, or visual data, Muse Spark is a strong contender.
Scientific Research and Reasoning
In Contemplating mode, Muse Spark scored 50.2% on Humanity’s Last Exam and 38.3% on FrontierScience Research, both ahead of GPT-5.4 Pro and Gemini Deep Think. These benchmarks test frontier-level scientific reasoning, and Muse Spark’s multi-agent approach gives it an edge on problems that benefit from parallel reasoning.
Where Meta Muse Spark Falls Short
Meta itself has acknowledged these gaps, describing them as areas of “continued investment.”
Coding and Software Development
This was the original release’s weakest area, and it is the one Meta has moved furthest on. At launch, a Terminal-Bench 2.0 score of 59.0 left Muse Spark well behind the coding leaders of the day. That gap has largely closed. Meta’s own launch chart for Muse Spark 1.2 puts it at 82.9% on Terminal-Bench 2.1, second only to Claude Opus 5 at 86.7%. Meta now ships Muse Code too, a terminal coding agent for macOS and Linux built on that model.
Second place is still second place, and the margin is wider on sustained work than on single tasks; Muse Spark 1.2 scores 59.3% on DeepSWE 1.1 against Opus 5’s 65.0%. Those are Meta’s own figures, and independent houses score it lower still. For heavy day-to-day coding, Claude vs ChatGPT remain the stronger options, and Claude Opus 5 leads the best AI models ranking for complex coding tasks.
Abstract Reasoning
The ARC-AGI-2 benchmark exposes the biggest gap. Muse Spark scores 42.5, while both GPT-5.4 (76.1) and Gemini 3.1 Pro (76.5) score nearly double. This benchmark tests novel pattern recognition and abstract problem-solving, requiring the model to identify visual patterns it has never seen before and generalize from minimal examples. The fact that Muse Spark scores less than half of its competitors here suggests the model’s architecture may not handle out-of-distribution reasoning as well as it handles knowledge-intensive tasks. For users who need AI for creative problem-solving or unusual analytical tasks, this is a meaningful limitation.
Agentic and Office Tasks
On GDPval-AA, which measures performance on real desktop and office tasks, the April release scored 1,444 ELO, well behind Claude Opus 4.6 (1,607) and GPT-5.4 (1,674). This benchmark tests whether an AI can autonomously complete multi-step workflows like filling out spreadsheets, navigating websites and managing documents. A low score means the model is less reliable when you hand it a long sequential task without supervision.
Meta named this a priority area, and agentic work is where the later releases improved most. Artificial Analysis has since rescored the line on a revised GDPval-AA v2, so those numbers are not directly comparable to the 1,444 above. Our Muse Spark 1.2 guide carries the current figures, plus the three different Terminal-Bench results the benchmark houses report for it.
Muse Spark Reasoning Modes Explained
One of Muse Spark’s most interesting features is its tiered approach to reasoning. Instead of a single processing mode, it offers three distinct levels.
Instant mode handles casual, everyday queries with fast response times. Think quick questions, simple lookups, and conversational exchanges. This is the default mode for most interactions.
Thinking mode adds deeper analysis when you need more thorough reasoning. The model takes extra processing time to work through complex problems step by step. This is comparable to the reasoning modes you find in ChatGPT vs Gemini and other frontier models.
Contemplating mode is where Muse Spark genuinely differentiates itself. Rather than a single model reasoning harder, it orchestrates multiple AI agents that reason in parallel and synthesize their findings. This multi-agent approach achieved 58% on Humanity’s Last Exam (with tools) and 38% on FrontierScience Research. Meta is rolling out Contemplating mode gradually, so it may not be available to all users immediately.
How to Use Meta Muse Spark
Getting started with Muse Spark is straightforward, and it is completely free. There is no waitlist, no subscription, and no account creation beyond your existing Meta login.
- Visit meta.ai in your browser or download the Meta AI app on your phone
- Start a conversation by typing a question or uploading an image
- Select your reasoning mode if available (Instant is the default; Thinking and Contemplating may require opt-in)
- Use multimodal features by photographing objects, sharing screenshots, or uploading images for analysis
- Try the shopping assistant to compare products, get pros and cons, and find purchase links
To get the most out of Muse Spark, match the reasoning mode to your task. Use Instant for quick factual lookups and casual conversation. Switch to Thinking when you need the model to work through a multi-step problem, analyze a document, or provide a detailed explanation. Reserve Contemplating for genuinely hard problems where you want multiple perspectives synthesized, like research questions or complex decisions. Keep in mind that Contemplating mode uses more processing time, so it is best saved for tasks where accuracy matters more than speed.
For visual tasks, Muse Spark works well when you upload images directly. You can photograph a product to get a detailed breakdown, share a chart for analysis, or snap a picture of a home appliance for troubleshooting guidance. The model generates annotated text responses that reference specific parts of your image.
Muse Spark also runs across Facebook, Instagram and WhatsApp. If you already use Meta AI on any of these platforms, you get the current Muse Spark model automatically, integrated directly into your existing chat interfaces.
For developers, Meta has since launched Muse Spark 1.1 and then Muse Spark 1.2 alongside Muse Code, through a public Meta Model API. Pricing has held at $1.25 per million input tokens and $4.25 per million output tokens across both upgrades, with $20 in free credits for new accounts. You can check the Meta AI blog for updates on API availability and open-source plans.
Meta Muse Spark vs. Other Free AI Options
If you regularly use AI assistants, you might be wondering how Muse Spark fits alongside models you can already access. The AI landscape now has several strong free or affordable options, and understanding where each model excels helps you pick the right tool for each task.
Muse Spark’s free tier is among the most generous of any frontier model, though it no longer stands alone since OpenAI moved free ChatGPT to GPT-5.6 Luna with unlimited text chats on August 6, 2026. There are no subscription fees, and Thinking mode is available to everyone with a Meta account. The tradeoff is that rate limits may apply for heavy users, and the free route keeps you inside Meta’s own apps; integrating the model into your own workflows means paying for the Meta Model API.
ChatGPT offers a free tier as well, running GPT-5.6 Luna since August 6, 2026, but the most powerful features require a $20/month Plus subscription. GPT-5.6 is the better choice for coding, agentic tasks, and abstract reasoning based on current benchmarks. Gemini 3.1 Pro has a generous free tier through Google AI Studio and is strong on multimodal work, though it sits behind both on the Intelligence Index at 48.
For users who want access to all major models without managing multiple subscriptions, Fello AI provides ChatGPT, Claude, Gemini, Grok, and DeepSeek in a single app for $9.99/month, rated 4.7 stars across 27,000+ reviews. This is particularly useful when you want to compare outputs across models or route specific tasks to the model that handles them best. You can ask the same question to multiple AI models and see which answer is most helpful for your specific use case.
The practical recommendation: use Muse Spark for health queries, visual analysis, and general conversation where it is free and competitive. Use other AI models for coding, complex reasoning, and productivity workflows where they have a measurable edge.
What Meta Muse Spark Means for the AI Landscape
Muse Spark represents a significant shift in Meta’s AI strategy. Rather than relying solely on the open-source Llama family, Meta now has a competitive closed model that can go head-to-head with the best from OpenAI, Google, and Anthropic. This builds on Meta’s broader push to compete directly with ChatGPT, Claude, and Grok through its own standalone AI products.
The model also validates the approach of building specialized strengths rather than chasing state-of-the-art scores everywhere. Muse Spark’s dominance in health AI and multimodal vision, combined with its token efficiency, shows that Meta is targeting specific high-value use cases where it can genuinely lead.
Meta’s decision to keep Muse Spark free also pressures competitors. While ChatGPT and Claude charge subscription fees for their most capable models, Meta is betting that free access will drive adoption across its 3+ billion user base on Facebook, Instagram, and WhatsApp. Whether Muse Spark’s open-source release materializes will shape how far that reach goes; on coding and agentic work, the gaps that defined the April release have already narrowed sharply across the 1.1 and 1.2 upgrades.
Conclusion
Meta Muse Spark is a genuinely capable AI model that excels in health reasoning, multimodal vision, and scientific research. It is free, accessible, and backed by Meta’s massive distribution network. Its launch weaknesses in coding and agentic tasks were real, and Meta’s rapid iteration has since narrowed them, most visibly with Muse Spark 1.2 and Muse Code. If you want to try Muse Spark, head to meta.ai today. For the best results across all AI tasks, consider using multiple models through Fello AI to match each task to the model that handles it best.
FAQ
What is Meta Muse Spark?
Meta Muse Spark is the first AI model from Meta Superintelligence Labs, built from the ground up with new infrastructure and architecture. It is a natively multimodal reasoning model that accepts text, image, and voice inputs.
Is Muse Spark free to use?
Yes. Muse Spark is completely free through meta.ai and the Meta AI app. Meta may impose rate limits for heavy usage, but there are no subscription fees or paywalls.
Is Muse Spark better than ChatGPT?
It depends on the task. At launch Muse Spark beat GPT-5.4 on health benchmarks (42.8 vs 40.1), scientific reasoning and chart understanding, while trailing badly on coding at 59.0 on Terminal-Bench 2.0. The coding gap has since narrowed a long way, and Meta’s own charts put Muse Spark 1.2 at 82.9% on Terminal-Bench 2.1. OpenAI’s current family is GPT-5.6, and neither model is better across the board.
Is Muse Spark open source?
Not currently. Unlike Meta’s Llama models, Muse Spark launched as a closed model. Meta has stated plans to release open-source weights in the future, but no specific timeline has been announced.
How is Muse Spark different from Llama?
Llama is Meta’s open-source model family available for developers to download and run locally. Muse Spark is a closed, consumer-facing model built by a separate team (Meta Superintelligence Labs) with a completely different architecture. Muse Spark is more capable but only accessible through Meta’s platforms.