Gemini 3.8 Flash vs Muse Spark 1.3: Benchmark Comparison
Gemini 3.8 Flash has the higher score on 8 of 14 shared benchmarks; Muse Spark 1.3 leads on 6.
The largest observed score gap is 25.66 pts on CyberBench v1.1 , where Muse Spark 1.3 leads.
Reported ±1 standard-error ranges overlap on 7 of 14 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Gemini 3.8 Flash | Muse Spark 1.3 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 54.83% ±1.03 | 53.20% ±1.07 | 1.63 pts | Reported ±1 SE ranges overlap |
| Harvey's Legal Agent Benchmark | 10.00% ±2.42 | 22.92% ±3.42 | 12.92 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 38.94% ±3.39 | 40.87% ±3.42 | 1.92 pts | Reported ±1 SE ranges overlap |
| EMB | 72.20% ±2.42 | 62.71% ±2.96 | 9.48 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 61.44% ±0.13 | 58.90% ±0.44 | 2.53 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 66.77% ±3.15 | 71.93% ±2.92 | 5.16 pts | Reported ±1 SE ranges overlap |
| ProofBench v1.1 | 48.00% ±5.02 | 55.00% ±5.00 | 7.00 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench Science | 8.57% ±3.37 | 4.29% ±2.44 | 4.29 pts | Reported ±1 SE ranges overlap |
| Code Migration | 36.55% ±4.18 | 27.58% ±4.07 | 8.97 pts | Reported ±1 SE ranges do not overlap |
| IOI | 56.94% ±2.26 | 43.94% ±1.45 | 13.00 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 1.00% ±0.70 | 0.50% ±0.50 | 0.50 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench 4.0 | 19.19% ±2.52 | 10.61% ±1.75 | 8.59 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 78.65% ±3.88 | 82.86% ±2.90 | 4.21 pts | Reported ±1 SE ranges overlap |
| CyberBench v1.1 | 43.75% ±2.21 | 69.41% ±5.76 | 25.66 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Gemini 3.8 Flash average | Muse Spark 1.3 average |
|---|---|---|
| Legal | 24.47% | 31.89% |
| Finance | 66.80% | 64.52% |
| Math | 48.00% | 55.00% |
| Science | 8.57% | 4.29% |
| Coding | 38.47% | 33.10% |
| Cyber | 43.75% | 69.41% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Gemini 3.8 Flash cost | Muse Spark 1.3 cost | Gemini 3.8 Flash latency | Muse Spark 1.3 latency |
|---|---|---|---|---|
| Vals Index | $5.73 | $3.38 | 51m57s | 32m42s |
| Harvey's Legal Agent Benchmark | $3.66 | $2.76 | 29m21s | 16m39s |
| Legal Research Bench | $1.63 | $0.65 | 5m24s | 8m51s |
| EMB | $8.24 | $3.55 | 12m49s | 22m55s |
| Finance Agent (v2) | $2.00 | $0.74 | 3m21s | 5m42s |
| Tax Agent Bench | $0.86 | $0.32 | 2m58s | 8m25s |
| ProofBench v1.1 | $0.60 | $0.43 | 6m09s | 8m54s |
| Terminal-Bench Science | $5.64 | $4.41 | 55m11s | 1h08m |
| Code Migration | $18.49 | $4.62 | 2h37m | 1h08m |
| IOI | $3.98 | $2.36 | 14m20s | 29m51s |
| ProgramBench | $10.93 | $11.56 | 46m39s | 1h09m |
| Terminal-Bench 4.0 | $8.77 | $7.04 | 1h48m | 1h59m |
| Vibe Code Bench v1.1 | $6.87 | $2.10 | 8m39s | 12m54s |
| CyberBench v1.1 | $1.31 | $4.26 | 9m49s | 24m53s |
Results available only for Gemini 3.8 Flash
- Vals RSI Index
- LegalBench
- MortgageTax
- TaxEval v2
- MedCode
- MedScribe
- BioMysteryBench
- MysteryMechanism
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- SAGE
- LiveCodeBench
- SkillsBench
- SWE-bench
- Vibe Code Bench 1-100
- Public Benefits Bench v1.1
- CUA-bench
- Time Horizon Index: KSP
Results available only for Muse Spark 1.3
None.