Jul 9, 2026
Meta's Muse Spark 1.1 evaluated across our benchmark suite
We evaluated Metaโs new Muse Spark 1.1 across our benchmark suite.
-
It debuts at #14 on the Vals Index (54.75%), narrowly behind GPT 5.5 (57.41%). At $1.25/test, it is the fourth-lowest-cost model in the top 20; it also averaged 827.8 seconds on the index, roughly 2โ4ร faster than the top three models.
-
It takes #2 on Finance Agent v2 (57.21%), just 0.65 points behind Gemini 3.5 Flash.
-
The largest generational gain is on Vibe Code Bench. Compared with Muse Spark, Muse Spark 1.1 rises from #43 to #5 and improves by 52.48 points (19.67% โ 72.16%).
-
Muse Spark 1.1 sets new highs on domain-specific work: #1 on MedScribe (88.89%), #1 on TaxEval v2 (79.72%), and #1 on Harveyโs Legal Agent Benchmark (20.00%). On MedScribe, it is nearly twice as fast as the #2 model.
The model has a 1M-token context window and supports up to 256k output tokens. Evaluations were run with reasoning effort set to โxhighโ and default temperature and top-p.
Congrats to the Meta team on the release!