Gemini 3.6 Flash vs Muse Spark 1.2: Benchmark Comparison

Gemini 3.6 Flash has the higher score on 9 of 21 shared benchmarks; Muse Spark 1.2 leads on 12.

The largest observed score gap is 24.17 pts on CyberBench v1.1 , where Muse Spark 1.2 leads.

Reported ±1 standard-error ranges overlap on 4 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.6 Flash Muse Spark 1.2 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 3.33% ±1.17 25.42% ±3.67 22.08 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 25.00% ±3.01 43.75% ±3.45 18.75 pts Reported ±1 SE ranges do not overlap
LegalBench 86.70% ±0.41 85.26% ±0.45 1.44 pts Reported ±1 SE ranges do not overlap
EMB 65.41% ±2.65 56.98% ±3.13 8.43 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 56.30% ±0.18 60.60% ±0.28 4.30 pts Reported ±1 SE ranges do not overlap
MortgageTax 67.61% ±0.92 65.42% ±0.93 2.19 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 49.90% ±3.30 56.86% ±2.32 6.96 pts Reported ±1 SE ranges do not overlap
TaxEval v2 74.86% ±0.85 80.38% ±0.76 5.52 pts Reported ±1 SE ranges do not overlap
MedCode 53.15% ±2.16 49.35% ±2.19 3.81 pts Reported ±1 SE ranges overlap
MedScribe 79.66% ±1.86 90.06% ±1.96 10.40 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 58.52% ±1.61 64.81% ±1.33 6.30 pts Reported ±1 SE ranges do not overlap
MMLU Pro 89.28% ±0.30 88.28% ±0.32 1.00 pts Reported ±1 SE ranges do not overlap
MMMU Pro 88.38% ±0.77 86.13% ±0.83 2.26 pts Reported ±1 SE ranges do not overlap
SAGE 48.69% ±3.38 47.66% ±3.44 1.03 pts Reported ±1 SE ranges overlap
Code Migration 30.93% ±4.07 29.95% ±4.02 0.97 pts Reported ±1 SE ranges overlap
IOI 35.06% ±5.70 21.78% ±0.87 13.28 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
SWE-bench 79.60% ±1.80 86.60% ±1.52 7.00 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 64.00% ±4.29 79.10% ±3.31 15.09 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 45.36% ±3.75 69.52% ±5.56 24.17 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 56.56% ±1.29 68.47% ±1.21 11.91 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.6 Flash average Muse Spark 1.2 average
Legal 38.35% 51.48%
Finance 62.81% 64.05%
Healthcare 66.41% 69.70%
Science 58.52% 64.81%
Academic 88.83% 87.20%
Education 48.69% 47.66%
Coding 41.92% 43.59%
Cyber 45.36% 69.52%
Social Mobility 56.56% 68.47%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.6 Flash cost Muse Spark 1.2 cost Gemini 3.6 Flash latency Muse Spark 1.2 latency
Harvey's Legal Agent Benchmark $1.90 $2.09 16m24s 24m16s
Legal Research Bench $0.47 $0.51 2m48s 5m07s
LegalBench N/A N/A 2.96s 30.10s
EMB $3.89 $2.31 12m06s 17m07s
Finance Agent (v2) $1.40 $0.77 3m45s 5m04s
MortgageTax N/A N/A 9.24s 44.62s
Tax Agent Bench $0.36 $0.27 105.32s 119.91s
TaxEval v2 N/A N/A 12.34s 48.31s
MedCode N/A N/A 16.68s 60.17s
MedScribe N/A N/A 33.55s 61.46s
BioMysteryBench $1.25 $0.99 12m42s 23m37s
MMLU Pro N/A N/A 7.30s 44.23s
MMMU Pro N/A N/A 11.31s 73.84s
SAGE N/A N/A 21.96s 29.07s
Code Migration $8.93 $3.78 1h30m 31m07s
IOI $8.67 $2.64 48m23s 25m35s
ProgramBench $5.92 $2.47 47m33s 26m38s
SWE-bench $1.19 $0.55 13m08s 8m07s
Vibe Code Bench v1.1 $3.04 $1.53 25m23s 19m55s
CyberBench v1.1 $1.31 $2.15 11m00s 20m21s
Public Benefits Bench v1.1 $0.39 $0.59 2m02s 8m55s

Results available only for Gemini 3.6 Flash

  • Terminal-Bench Science
  • GPQA Diamond
  • LiveCodeBench

Results available only for Muse Spark 1.2

  • Vals Index
  • ProofBench v1.1
  • SkillsBench
  • Terminal-Bench 4.0
Model details Gemini 3.6 Flash Model details Muse Spark 1.2