Stay connected with KayaToday, follow us on Instagram and Facebook for the latest news and reviews delivered straight to you.
There is a reliable pattern in frontier AI launches: the model that ships is not always the model that headlines the announcement. Meta’s release of Muse Spark 1.3 on Wednesday follows that pattern closely enough to warrant a careful read before any enterprise team adjusts its stack.
Meta co-founder and CEO Mark Zuckerberg called it the company’s “biggest jump” yet in coding and agentic work, writing on X that Muse Spark 1.3 offers “frontier performance almost too cheap to meter.” Both halves of that sentence contain real substance, and both contain meaningful qualifications that the launch materials do not emphasise.
The Model You Can Use Today Is Not the One Setting the Benchmarks
The strongest benchmark numbers Meta published belong to Muse Spark 1.3’s “max” reasoning configuration. That version is still completing safety testing and will arrive, in Meta’s words, “shortly.” Independent benchmarking firm Artificial Analysis evaluated max in a limited partner preview and currently lists no API provider for that configuration at all, meaning no developer can broadly access it today.
What is actually rolling out this week through Meta’s Muse Code harness and the Meta Model API uses the previously available reasoning settings, including the “xhigh” tier. Meta does disclose results for both configurations in its underlying evaluation report, so the company is not hiding the gap. But its launch materials prominently feature max scores, and some of the largest figures belong to that version.
The differences are real but not dramatic. Meta reports GDPval-AA v2 scores of 1,754 Elo for max versus 1,709 for xhigh, OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2. On some tests the distinction disappears or reverses: DeepSearchQA is tied at 89.4, while xhigh scores 89.2 on Terminal-Bench 2.1 against max’s 88.8.
Artificial Analysis places Muse Spark 1.3 max at 62 on its Intelligence Index and the shipping xhigh version at 61. That 61 ties GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. Anthropic still leads the leaderboard: Claude Fable 5.1 reaches 66 at max and 65 at xhigh, while Claude Opus 5 reaches 63 at both settings. Muse Spark 1.3 xhigh is legitimately inside the frontier cluster. It is not currently the model defining that frontier.
That said, the gap from Muse Spark 1.2 is genuine. Last month’s release trailed Anthropic’s best model on most coding comparisons Meta presented, scoring 82.9% on Terminal-Bench 2.1 against Opus 5’s 86.7%. With 1.3, Meta is trading wins with both OpenAI and Anthropic across several coding and agentic evaluations rather than simply appearing in the contest.
What “Almost Too Cheap to Meter” Actually Means
Zuckerberg’s pricing claim is not about a price cut. Meta kept Standard API pricing identical to Muse Spark 1.2: $1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens. The argument is instead about what developers can accomplish per dollar spent.
Artificial Analysis provides partial support for that argument. It measures Muse Spark 1.3 xhigh at 235.2 output tokens per second and estimates a cost of $0.55 per Intelligence Index task, which at a score of 61 gives it the lowest cost per task of any currently measured model at that intelligence level. That is a meaningful efficiency claim.
The complication is that Muse Spark 1.2 cost only $0.40 per Artificial Analysis task while scoring 57. Despite unchanged per-token pricing, the cost of completing an average benchmark task therefore rose generation over generation. Artificial Analysis attributes this primarily to heavier input-token consumption on agentic evaluations. Meta separately reports that Muse Spark 1.3 used roughly 20% fewer tool calls and 25% fewer tokens than 1.2 in its own internal coding workflows, so the two figures are measuring different things rather than contradicting each other directly. But the episode illustrates why “cheap” becomes slippery once models operate as agents: token rates, reasoning effort, turn count, tool calls, and retries all feed into the actual cost of finishing a task.
Meta also retains its Contributor tier at $0.10 per million input tokens and $0.20 per million output tokens, in exchange for permission to use prompts and completions for model training. For prototyping that is attractive. For enterprises working with proprietary code or sensitive internal data, it represents a materially different data-governance calculation that deserves deliberate scrutiny before adoption.
Meta Versus Google: A Two-Point Lead and a Faster Rival
Meta chief AI officer Alexandr Wang added his own commentary when Artificial Analysis posted results, reposting them on X with the line: “i really hate to say it, but… gemini who?” The timing was pointed because Google released Gemini 3.8 Flash on the same day, targeting almost exactly the same workload category: long-horizon software engineering, autonomous agents, and multi-step professional reasoning.
Independent numbers give Wang something to work with, though not a decisive victory. Artificial Analysis scores Muse Spark 1.3 xhigh at 61 and $0.55 per task, against 59 and $0.58 for Gemini 3.8 Flash at high reasoning. Meta edges Google on both intelligence and task cost at those settings.
Google wins clearly on throughput. Artificial Analysis measures Gemini 3.8 Flash high at roughly 305 output tokens per second against Muse Spark’s 235, a gap of about 30%. Google also has the lower raw API sticker price for now, charging an introductory $0.75 per million input tokens and $3.75 per million output tokens. That promotional pricing expires December 31, 2026, after which it rises to $1.50 per million input tokens and $7.50 per million output tokens, putting it well above Meta’s Standard rate.
For an enterprise architect weighing the two today, the honest summary is that Muse Spark 1.3 xhigh is the slightly stronger high-effort agent by this independent measure, while Gemini 3.8 Flash is substantially faster and cheaper on raw tokens during its launch window. Wang’s trash talk is fun. The actual decision depends on whether throughput or reasoning depth matters more to a given workload.
The Open-Weight Question That Matters More Than Any Benchmark
For developers in Southeast Asia who built on Meta’s Llama family precisely because downloadable weights enabled self-hosting, fine-tuning, and control over inference costs, the more consequential issue may be Meta’s evolving stance on open weights rather than any leaderboard position.
When Meta launched Muse Code and Muse Spark 1.2 in August, the move to proprietary API-only products was a striking departure from the company’s years of arguing that open AI was the correct path forward. Five days after that launch, Meta partially reversed course: it released the 30-billion-parameter Muse Glimmer under an Apache 2.0 licence, and Zuckerberg stated that Spark 1.2 weights would follow “in the coming weeks.” Reuters separately reported the same plan.
Meta has now shipped Muse Spark 1.3 as another proprietary model. The new announcement no longer references Spark 1.2 weights specifically. It describes the roadmap as including “the Muse Spark open weights release” without specifying a version, date, model size, or licence. Zuckerberg’s X post similarly referred to “Muse Spark open weights releases” coming soon, plural and unscheduled.
That ambiguity does not constitute a broken promise, and “coming weeks” can reasonably stretch beyond three weeks. But the roadmap has become less clear with each successive announcement rather than more so.
Muse Spark 1.3 demonstrates that Meta can iterate proprietary frontier models at extraordinary speed. The shipping xhigh configuration is fast, competitively priced, and meaningfully closer to the top of independent rankings than any previous Muse Spark release. The max preview shows the family can push further still when given more reasoning compute. The next test is less about benchmarks and more about trust: whether Meta can deliver a roadmap that enterprises and open-source developers can actually plan around, including making its strongest capabilities broadly deployable and honouring the open-weight commitment it has already made publicly.