Release Date: Jul 9, 2026

Developer Metaย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 256k
Token Costs (in/out) $1.25/4.25
Weights Private
Input Modalities

Accuracy

54.75 % ยฑ 1.15

Cost / Test (Vals Index)

$ 1.251

Latency

13 min 48 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.15
15/46

0.0%

ยฑ4.07
18/49

0.0%

ยฑ3.15
22/46

0.0%

ยฑ0.79
5/49

0.0%

ยฑ0.00
7/7

0.0%

ยฑ3.37
15/49

0.0%

ยฑ1.95
3/84

0.0%

ยฑ0.93
36/95

0.0%

ยฑ3.28
29/76

0.0%

ยฑ1.27
15/30

0.0%

ยฑ0.77
2/140

0.0%

ยฑ3.33
13/84

0.0%

ยฑ2.10
21/133

0.0%

ยฑ0.99
27/138

0.0%

ยฑ0.42
22/137

0.0%

ยฑ0.32
15/133

0.0%

ยฑ0.82
16/89

0.0%

ยฑ0.00
22/39

0.0%

ยฑ4.44
9/28

0.0%

ยฑ1.72
17/83

0.0%

ยฑ0.99
17/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Meta
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 256,000
Reasoning Effort: xhigh

Updates

Jul 9, 2026

We evaluated Metaโ€™s new Muse Spark 1.1 across our benchmark suite.

  • It debuts at #14 on the Vals Index (54.75%), narrowly behind GPT 5.5 (57.41%). At $1.25/test, it is the fourth-lowest-cost model in the top 20; it also averaged 827.8 seconds on the index, roughly 2โ€“4ร— faster than the top three models.

  • It takes #2 on Finance Agent v2 (57.21%), just 0.65 points behind Gemini 3.5 Flash.

  • The largest generational gain is on Vibe Code Bench. Compared with Muse Spark, Muse Spark 1.1 rises from #43 to #5 and improves by 52.48 points (19.67% โ†’ 72.16%).

  • Muse Spark 1.1 sets new highs on domain-specific work: #1 on MedScribe (88.89%), #1 on TaxEval v2 (79.72%), and #1 on Harveyโ€™s Legal Agent Benchmark (20.00%). On MedScribe, it is nearly twice as fast as the #2 model.

The model has a 1M-token context window and supports up to 256k output tokens. Evaluations were run with reasoning effort set to โ€œxhighโ€ and default temperature and top-p.

Congrats to the Meta team on the release!