GPT-5.6 Luna vs MiMo V2.6 Flash: Benchmark Comparison
GPT-5.6 Luna has the higher score on 6 of 20 shared benchmarks; MiMo V2.6 Flash leads on 14.
The largest observed score gap is 14.11 pts on IOI , where GPT-5.6 Luna leads.
Reported ±1 standard-error ranges overlap on 13 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | GPT-5.6 Luna | MiMo V2.6 Flash | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 51.69% ±1.06 | 53.23% ±1.10 | 1.54 pts | Reported ±1 SE ranges overlap |
| Harvey's Legal Agent Benchmark | 1.25% ±0.83 | 11.25% ±2.54 | 10.00 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 36.54% ±3.35 | 37.98% ±3.37 | 1.44 pts | Reported ±1 SE ranges overlap |
| EMB | 67.12% ±2.94 | 65.46% ±2.71 | 1.66 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 55.04% ±0.31 | 56.28% ±0.46 | 1.23 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 60.81% ±3.26 | 59.90% ±3.13 | 0.91 pts | Reported ±1 SE ranges overlap |
| MedCode | 42.39% ±2.27 | 41.06% ±2.03 | 1.33 pts | Reported ±1 SE ranges overlap |
| MedScribe | 84.39% ±2.58 | 85.28% ±1.99 | 0.88 pts | Reported ±1 SE ranges overlap |
| ProofBench v1.1 | 60.00% ±4.92 | 63.00% ±4.85 | 3.00 pts | Reported ±1 SE ranges overlap |
| BioMysteryBench | 61.48% ±0.74 | 69.26% ±0.98 | 7.78 pts | Reported ±1 SE ranges do not overlap |
| MysteryMechanism | 14.41% ±2.36 | 21.62% ±2.77 | 7.21 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 0.00% ±0.00 | 4.29% ±2.44 | 4.29 pts | Reported ±1 SE ranges do not overlap |
| SAGE | 44.22% ±3.33 | 43.53% ±3.43 | 0.69 pts | Reported ±1 SE ranges overlap |
| Code Migration | 44.55% ±4.24 | 40.93% ±4.35 | 3.62 pts | Reported ±1 SE ranges overlap |
| IOI | 61.78% ±11.61 | 47.67% ±4.79 | 14.11 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 0.00% ±0.00 | 0.50% ±0.50 | 0.50 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench 4.0 | 11.62% ±1.01 | 24.24% ±1.51 | 12.63 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 77.06% ±3.07 | 78.96% ±4.04 | 1.90 pts | Reported ±1 SE ranges overlap |
| CyberBench v1.1 | 73.63% ±5.56 | 75.36% ±5.42 | 1.73 pts | Reported ±1 SE ranges overlap |
| Public Benefits Bench v1.1 | 61.16% ±1.27 | 67.59% ±1.22 | 6.43 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | GPT-5.6 Luna average | MiMo V2.6 Flash average |
|---|---|---|
| Legal | 18.89% | 24.62% |
| Finance | 60.99% | 60.55% |
| Healthcare | 63.39% | 63.17% |
| Math | 60.00% | 63.00% |
| Science | 25.30% | 31.72% |
| Education | 44.22% | 43.53% |
| Coding | 39.00% | 38.46% |
| Cyber | 73.63% | 75.36% |
| Social Mobility | 61.16% | 67.59% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | GPT-5.6 Luna cost | MiMo V2.6 Flash cost | GPT-5.6 Luna latency | MiMo V2.6 Flash latency |
|---|---|---|---|---|
| Vals Index | $0.82 | $0.20 | 29m27s | 58m02s |
| Harvey's Legal Agent Benchmark | $0.38 | $0.09 | 10m56s | 17m12s |
| Legal Research Bench | $0.85 | $0.07 | 39m34s | 15m20s |
| EMB | $0.37 | $0.11 | 11m00s | 25m22s |
| Finance Agent (v2) | $0.28 | $0.07 | 12m52s | 7m02s |
| Tax Agent Bench | $0.47 | $0.05 | 42m46s | 17m19s |
| MedCode | N/A | N/A | 81.28s | 93.37s |
| MedScribe | N/A | N/A | 116.89s | 74.48s |
| ProofBench v1.1 | $0.12 | $0.12 | 8m34s | 1h03m |
| BioMysteryBench | $0.09 | $0.05 | 18m31s | 28m28s |
| MysteryMechanism | $0.11 | $0.06 | 8m13s | 35m08s |
| Terminal-Bench Science | $0.58 | $0.24 | 2h03m | 3h59m |
| SAGE | N/A | N/A | 51.88s | 2m16s |
| Code Migration | $1.88 | $0.49 | 58m35s | 3h12m |
| IOI | $0.58 | $0.22 | 1h10m | 1h49m |
| ProgramBench | $0.94 | $0.47 | 41m34s | 3h22m |
| Terminal-Bench 4.0 | $0.73 | $0.22 | 35m55s | 2h25m |
| Vibe Code Bench v1.1 | $0.73 | $0.56 | 25m52s | 50m13s |
| CyberBench v1.1 | $0.40 | $0.05 | 16m15s | 28m22s |
| Public Benefits Bench v1.1 | $0.27 | $0.03 | 38m40s | 20m56s |
Results available only for GPT-5.6 Luna
- LegalBench
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- SkillsBench
- SWE-bench
- Vibe Code Bench 1-100
Results available only for MiMo V2.6 Flash
- SRE Bench