Gemini 3.8 Flash vs Muse Spark 1.3 Max: Benchmark Comparison

Gemini 3.8 Flash has the higher score on 5 of 18 shared benchmarks; Muse Spark 1.3 Max leads on 13.

The largest observed score gap is 28.99 pts on CyberBench v1.1 , where Muse Spark 1.3 Max leads.

Reported ±1 standard-error ranges overlap on 8 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.8 Flash Muse Spark 1.3 Max Gap Reported uncertainty
Vals Index 54.83% ±1.03 58.16% ±1.19 3.34 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 19.93% 19.64% 0.29 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 10.00% ±2.42 23.75% ±3.55 13.75 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 38.94% ±3.39 55.29% ±3.46 16.35 pts Reported ±1 SE ranges do not overlap
EMB 72.20% ±2.42 67.43% ±3.06 4.77 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 61.44% ±0.13 59.96% ±2.06 1.48 pts Reported ±1 SE ranges overlap
Tax Agent Bench 66.77% ±3.15 72.44% ±2.88 5.67 pts Reported ±1 SE ranges overlap
ProofBench v1.1 48.00% ±5.02 58.00% ±4.96 10.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 36.49% ±3.24 36.04% ±3.23 0.45 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 8.57% ±3.37 10.00% ±3.61 1.43 pts Reported ±1 SE ranges overlap
Code Migration 36.55% ±4.18 47.41% ±4.26 10.87 pts Reported ±1 SE ranges do not overlap
IOI 56.94% ±2.26 56.56% ±2.52 0.39 pts Reported ±1 SE ranges overlap
ProgramBench 1.00% ±0.70 2.50% ±1.11 1.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 19.19% ±2.52 24.75% ±0.51 5.55 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 18.77% ±3.31 20.46% ±3.85 1.69 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 78.65% ±3.88 85.86% ±2.51 7.20 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 43.75% ±2.21 72.74% ±5.67 28.99 pts Reported ±1 SE ranges do not overlap
CUA-bench 4.17% 5.83% 1.67 pts Uncertainty comparison unavailable

Performance by category

Category Gemini 3.8 Flash average Muse Spark 1.3 Max average
Legal 24.47% 39.52%
Finance 66.80% 66.61%
Math 48.00% 58.00%
Science 22.53% 23.02%
Coding 35.18% 39.59%
Cyber 43.75% 72.74%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.8 Flash cost Muse Spark 1.3 Max cost Gemini 3.8 Flash latency Muse Spark 1.3 Max latency
Vals Index $5.73 $3.79 51m57s 23m33s
Vals RSI Index $332.04 $48.82 90h00m 90h00m
Harvey's Legal Agent Benchmark $3.66 $2.26 29m21s 14m35s
Legal Research Bench $1.63 $0.59 5m24s 5m17s
EMB $8.24 $2.55 12m49s 13m50s
Finance Agent (v2) $2.00 $0.76 3m21s 3m19s
Tax Agent Bench $0.86 $0.39 2m58s 4m09s
ProofBench v1.1 $0.60 $0.48 6m09s 9m29s
MysteryMechanism $1.73 $0.85 5m22s 7m32s
Terminal-Bench Science $5.64 $6.09 55m11s 1h24m
Code Migration $18.49 $14.21 2h37m 1h24m
IOI $3.98 $1.73 14m20s 28m58s
ProgramBench $10.93 $54.29 46m39s 5h28m
Terminal-Bench 4.0 $8.77 $6.65 1h48m 52m56s
Vibe Code Bench 1-100 $18.91 $5.60 1h06m 20m02s
Vibe Code Bench v1.1 $6.87 $2.54 8m39s 16m18s
CyberBench v1.1 $1.31 $3.34 9m49s 16m18s
CUA-bench $280.03 $143.48 36.55s 36.39s

Results available only for Gemini 3.8 Flash

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • BioMysteryBench
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Public Benefits Bench v1.1
  • Time Horizon Index: KSP

Results available only for Muse Spark 1.3 Max

None.

Model details Gemini 3.8 Flash Model details Muse Spark 1.3 Max