GPT 5.5 vs MiMo V2.6 Pro: Benchmark Comparison
GPT 5.5 has the higher score on 5 of 13 shared benchmarks; MiMo V2.6 Pro leads on 7.
The largest observed score gap is 15.38 pts on Vibe Code Bench v1.1 , where MiMo V2.6 Pro leads.
Reported ±1 standard-error ranges overlap on 9 of 13 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | GPT 5.5 | MiMo V2.6 Pro | Gap | Reported uncertainty |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | 3.75% ±1.17 | 10.83% ±2.45 | 7.08 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 40.38% ±3.41 | 47.12% ±3.47 | 6.73 pts | Reported ±1 SE ranges overlap |
| EMB | 64.54% ±2.87 | 62.86% ±3.14 | 1.69 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 51.76% ±0.55 | 57.34% ±0.57 | 5.58 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 60.46% ±3.22 | 64.94% ±3.21 | 4.48 pts | Reported ±1 SE ranges overlap |
| MedCode | 49.10% ±2.19 | 44.97% ±2.10 | 4.13 pts | Reported ±1 SE ranges overlap |
| MedScribe | 86.87% ±1.93 | 88.31% ±1.94 | 1.44 pts | Reported ±1 SE ranges overlap |
| SAGE | 51.53% ±3.95 | 45.05% ±3.40 | 6.48 pts | Reported ±1 SE ranges overlap |
| Code Migration | 45.16% ±4.16 | 43.01% ±4.32 | 2.15 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 0.50% ±0.50 | 0.50% ±0.50 | 0.00 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 69.85% ±4.54 | 85.22% ±3.39 | 15.38 pts | Reported ±1 SE ranges do not overlap |
| SRE Bench | 3.82% ±1.19 | 3.05% ±1.06 | 0.76 pts | Reported ±1 SE ranges overlap |
| Public Benefits Bench v1.1 | 60.89% ±1.27 | 68.94% ±1.20 | 8.05 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | GPT 5.5 average | MiMo V2.6 Pro average |
|---|---|---|
| Legal | 22.07% | 28.97% |
| Finance | 58.92% | 61.71% |
| Healthcare | 67.98% | 66.64% |
| Education | 51.53% | 45.05% |
| Coding | 38.50% | 42.91% |
| Cyber | 3.82% | 3.05% |
| Social Mobility | 60.89% | 68.94% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | GPT 5.5 cost | MiMo V2.6 Pro cost | GPT 5.5 latency | MiMo V2.6 Pro latency |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | $4.60 | $0.22 | 12m14s | 25m49s |
| Legal Research Bench | $7.40 | $0.18 | 37m45s | 30m16s |
| EMB | $3.27 | $0.33 | 15m11s | 55m14s |
| Finance Agent (v2) | $4.15 | $0.20 | 11m02s | 10m48s |
| Tax Agent Bench | $4.07 | $0.13 | 15m52s | 25m48s |
| MedCode | N/A | N/A | 2m40s | 2m57s |
| MedScribe | N/A | N/A | 2m13s | 2m53s |
| SAGE | N/A | N/A | 76.05s | 3m40s |
| Code Migration | $6.44 | $0.69 | 31m35s | 2h49m |
| ProgramBench | $6.95 | $0.73 | 22m03s | 3h00m |
| Vibe Code Bench v1.1 | $16.66 | $1.04 | 31m52s | 1h01m |
| SRE Bench | $17.77 | $0.53 | 46m28s | 2h23m |
| Public Benefits Bench v1.1 | $3.97 | $0.08 | 39m37s | 37m14s |
Results available only for GPT 5.5
- Vals RSI Index
- LegalBench
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- LiveCodeBench
- SkillsBench
- SWE-bench
- Time Horizon Index: KSP
Results available only for MiMo V2.6 Pro
- Vals Index
- ProofBench v1.1
- MysteryMechanism
- Terminal-Bench Science
- IOI
- Terminal-Bench 4.0
- CyberBench v1.1