Grok 4.7 vs Muse Spark 1.3: Benchmark Comparison

Grok 4.7 has the higher score on 9 of 14 shared benchmarks; Muse Spark 1.3 leads on 4.

The largest observed score gap is 29.00 pts on ProofBench v1.1 , where Muse Spark 1.3 leads.

Reported ±1 standard-error ranges overlap on 7 of 14 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.7 Muse Spark 1.3 Gap Reported uncertainty
Vals Index 54.95% ±1.07 53.20% ±1.07 1.75 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 12.50% ±2.66 22.92% ±3.42 10.42 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 47.12% ±3.47 40.87% ±3.42 6.25 pts Reported ±1 SE ranges overlap
EMB 66.99% ±3.04 62.71% ±2.96 4.28 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 52.25% ±0.35 58.90% ±0.44 6.65 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 65.60% ±3.24 71.93% ±2.92 6.34 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 26.00% ±4.41 55.00% ±5.00 29.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 10.00% ±3.61 4.29% ±2.44 5.71 pts Reported ±1 SE ranges overlap
Code Migration 44.82% ±4.21 27.58% ±4.07 17.24 pts Reported ±1 SE ranges do not overlap
IOI 57.72% ±1.99 43.94% ±1.45 13.78 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 28.79% ±2.31 10.61% ±1.75 18.18 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 86.17% ±2.18 82.86% ±2.90 3.32 pts Reported ±1 SE ranges overlap
CyberBench v1.1 69.46% ±5.67 69.41% ±5.76 0.06 pts Reported ±1 SE ranges overlap

Performance by category

Category Grok 4.7 average Muse Spark 1.3 average
Legal 29.81% 31.89%
Finance 61.61% 64.52%
Math 26.00% 55.00%
Science 10.00% 4.29%
Coding 43.60% 33.10%
Cyber 69.46% 69.41%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.7 cost Muse Spark 1.3 cost Grok 4.7 latency Muse Spark 1.3 latency
Vals Index $12.12 $3.38 35m31s 32m42s
Harvey's Legal Agent Benchmark $11.13 $2.76 42m05s 16m39s
Legal Research Bench $4.92 $0.65 21m14s 8m51s
EMB $6.48 $3.55 34m57s 22m55s
Finance Agent (v2) $2.72 $0.74 15m33s 5m42s
Tax Agent Bench $1.80 $0.32 16m36s 8m25s
ProofBench v1.1 $0.79 $0.43 12m47s 8m54s
Terminal-Bench Science $14.45 $4.41 1h29m 1h08m
Code Migration $36.55 $4.62 1h09m 1h08m
IOI $12.71 $2.36 46m48s 29m51s
ProgramBench $46.49 $11.56 3h38m 1h09m
Terminal-Bench 4.0 $18.09 $7.04 50m21s 1h59m
Vibe Code Bench v1.1 $15.83 $2.10 35m47s 12m54s
CyberBench v1.1 $8.54 $4.26 15m30s 24m53s

Results available only for Grok 4.7

  • Vals RSI Index
  • LegalBench
  • MedCode
  • MedScribe
  • BioMysteryBench
  • MysteryMechanism
  • SAGE
  • Public Benefits Bench v1.1

Results available only for Muse Spark 1.3

None.

Model details Grok 4.7 Model details Muse Spark 1.3