GPT-5.6 Luna vs MiMo V2.6 Pro: Benchmark Comparison
GPT-5.6 Luna has the higher score on 4 of 19 shared benchmarks; MiMo V2.6 Pro leads on 15.
The largest observed score gap is 22.45 pts on IOI , where GPT-5.6 Luna leads.
Reported ±1 standard-error ranges overlap on 9 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | GPT-5.6 Luna | MiMo V2.6 Pro | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 51.69% ±1.06 | 55.20% ±1.18 | 3.51 pts | Reported ±1 SE ranges do not overlap |
| Harvey's Legal Agent Benchmark | 1.25% ±0.83 | 10.83% ±2.45 | 9.58 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 36.54% ±3.35 | 47.12% ±3.47 | 10.58 pts | Reported ±1 SE ranges do not overlap |
| EMB | 67.12% ±2.94 | 62.86% ±3.14 | 4.26 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 55.04% ±0.31 | 57.34% ±0.57 | 2.30 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 60.81% ±3.26 | 64.94% ±3.21 | 4.12 pts | Reported ±1 SE ranges overlap |
| MedCode | 42.39% ±2.27 | 44.97% ±2.10 | 2.58 pts | Reported ±1 SE ranges overlap |
| MedScribe | 84.39% ±2.58 | 88.31% ±1.94 | 3.92 pts | Reported ±1 SE ranges overlap |
| ProofBench v1.1 | 60.00% ±4.92 | 70.00% ±4.61 | 10.00 pts | Reported ±1 SE ranges do not overlap |
| MysteryMechanism | 14.41% ±2.36 | 15.31% ±2.42 | 0.90 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench Science | 0.00% ±0.00 | 2.86% ±2.01 | 2.86 pts | Reported ±1 SE ranges do not overlap |
| SAGE | 44.22% ±3.33 | 45.05% ±3.40 | 0.83 pts | Reported ±1 SE ranges overlap |
| Code Migration | 44.55% ±4.24 | 43.01% ±4.32 | 1.54 pts | Reported ±1 SE ranges overlap |
| IOI | 61.78% ±11.61 | 39.33% ±2.41 | 22.45 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 0.00% ±0.00 | 0.50% ±0.50 | 0.50 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench 4.0 | 11.62% ±1.01 | 31.31% ±3.07 | 19.70 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 77.06% ±3.07 | 85.22% ±3.39 | 8.17 pts | Reported ±1 SE ranges do not overlap |
| CyberBench v1.1 | 73.63% ±5.56 | 72.86% ±5.50 | 0.77 pts | Reported ±1 SE ranges overlap |
| Public Benefits Bench v1.1 | 61.16% ±1.27 | 68.94% ±1.20 | 7.78 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | GPT-5.6 Luna average | MiMo V2.6 Pro average |
|---|---|---|
| Legal | 18.89% | 28.97% |
| Finance | 60.99% | 61.71% |
| Healthcare | 63.39% | 66.64% |
| Math | 60.00% | 70.00% |
| Science | 7.21% | 9.09% |
| Education | 44.22% | 45.05% |
| Coding | 39.00% | 39.88% |
| Cyber | 73.63% | 72.86% |
| Social Mobility | 61.16% | 68.94% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | GPT-5.6 Luna cost | MiMo V2.6 Pro cost | GPT-5.6 Luna latency | MiMo V2.6 Pro latency |
|---|---|---|---|---|
| Vals Index | $0.82 | $0.41 | 29m27s | 1h07m |
| Harvey's Legal Agent Benchmark | $0.38 | $0.22 | 10m56s | 25m49s |
| Legal Research Bench | $0.85 | $0.18 | 39m34s | 30m16s |
| EMB | $0.37 | $0.33 | 11m00s | 55m14s |
| Finance Agent (v2) | $0.28 | $0.20 | 12m52s | 10m48s |
| Tax Agent Bench | $0.47 | $0.13 | 42m46s | 25m48s |
| MedCode | N/A | N/A | 81.28s | 2m57s |
| MedScribe | N/A | N/A | 116.89s | 2m53s |
| ProofBench v1.1 | $0.12 | $0.25 | 8m34s | 1h02m |
| MysteryMechanism | $0.11 | $0.15 | 8m13s | 44m00s |
| Terminal-Bench Science | $0.58 | $0.62 | 2h03m | 3h12m |
| SAGE | N/A | N/A | 51.88s | 3m40s |
| Code Migration | $1.88 | $0.69 | 58m35s | 2h49m |
| IOI | $0.58 | $0.84 | 1h10m | 2h26m |
| ProgramBench | $0.94 | $0.73 | 41m34s | 3h00m |
| Terminal-Bench 4.0 | $0.73 | $0.50 | 35m55s | 2h42m |
| Vibe Code Bench v1.1 | $0.73 | $1.04 | 25m52s | 1h01m |
| CyberBench v1.1 | $0.40 | $0.09 | 16m15s | 27m41s |
| Public Benefits Bench v1.1 | $0.27 | $0.08 | 38m40s | 37m14s |
Results available only for GPT-5.6 Luna
- LegalBench
- MortgageTax
- TaxEval v2
- BioMysteryBench
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- SkillsBench
- SWE-bench
- Vibe Code Bench 1-100
Results available only for MiMo V2.6 Pro
- SRE Bench