Grok 4.7 vs Muse Spark 1.3: Benchmark Comparison
Grok 4.7 has the higher score on 9 of 14 shared benchmarks; Muse Spark 1.3 leads on 4.
The largest observed score gap is 29.00 pts on ProofBench v1.1 , where Muse Spark 1.3 leads.
Reported ±1 standard-error ranges overlap on 7 of 14 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Grok 4.7 | Muse Spark 1.3 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 54.95% ±1.07 | 53.20% ±1.07 | 1.75 pts | Reported ±1 SE ranges overlap |
| Harvey's Legal Agent Benchmark | 12.50% ±2.66 | 22.92% ±3.42 | 10.42 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 47.12% ±3.47 | 40.87% ±3.42 | 6.25 pts | Reported ±1 SE ranges overlap |
| EMB | 66.99% ±3.04 | 62.71% ±2.96 | 4.28 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 52.25% ±0.35 | 58.90% ±0.44 | 6.65 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 65.60% ±3.24 | 71.93% ±2.92 | 6.34 pts | Reported ±1 SE ranges do not overlap |
| ProofBench v1.1 | 26.00% ±4.41 | 55.00% ±5.00 | 29.00 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 10.00% ±3.61 | 4.29% ±2.44 | 5.71 pts | Reported ±1 SE ranges overlap |
| Code Migration | 44.82% ±4.21 | 27.58% ±4.07 | 17.24 pts | Reported ±1 SE ranges do not overlap |
| IOI | 57.72% ±1.99 | 43.94% ±1.45 | 13.78 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 0.50% ±0.50 | 0.50% ±0.50 | 0.00 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench 4.0 | 28.79% ±2.31 | 10.61% ±1.75 | 18.18 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 86.17% ±2.18 | 82.86% ±2.90 | 3.32 pts | Reported ±1 SE ranges overlap |
| CyberBench v1.1 | 69.46% ±5.67 | 69.41% ±5.76 | 0.06 pts | Reported ±1 SE ranges overlap |
Performance by category
| Category | Grok 4.7 average | Muse Spark 1.3 average |
|---|---|---|
| Legal | 29.81% | 31.89% |
| Finance | 61.61% | 64.52% |
| Math | 26.00% | 55.00% |
| Science | 10.00% | 4.29% |
| Coding | 43.60% | 33.10% |
| Cyber | 69.46% | 69.41% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Grok 4.7 cost | Muse Spark 1.3 cost | Grok 4.7 latency | Muse Spark 1.3 latency |
|---|---|---|---|---|
| Vals Index | $12.12 | $3.38 | 35m31s | 32m42s |
| Harvey's Legal Agent Benchmark | $11.13 | $2.76 | 42m05s | 16m39s |
| Legal Research Bench | $4.92 | $0.65 | 21m14s | 8m51s |
| EMB | $6.48 | $3.55 | 34m57s | 22m55s |
| Finance Agent (v2) | $2.72 | $0.74 | 15m33s | 5m42s |
| Tax Agent Bench | $1.80 | $0.32 | 16m36s | 8m25s |
| ProofBench v1.1 | $0.79 | $0.43 | 12m47s | 8m54s |
| Terminal-Bench Science | $14.45 | $4.41 | 1h29m | 1h08m |
| Code Migration | $36.55 | $4.62 | 1h09m | 1h08m |
| IOI | $12.71 | $2.36 | 46m48s | 29m51s |
| ProgramBench | $46.49 | $11.56 | 3h38m | 1h09m |
| Terminal-Bench 4.0 | $18.09 | $7.04 | 50m21s | 1h59m |
| Vibe Code Bench v1.1 | $15.83 | $2.10 | 35m47s | 12m54s |
| CyberBench v1.1 | $8.54 | $4.26 | 15m30s | 24m53s |
Results available only for Grok 4.7
- Vals RSI Index
- LegalBench
- MedCode
- MedScribe
- BioMysteryBench
- MysteryMechanism
- SAGE
- Public Benefits Bench v1.1
Results available only for Muse Spark 1.3
None.