GPT-5.6 Luna vs MiMo V2.6 Pro: Benchmark Comparison

GPT-5.6 Luna has the higher score on 4 of 19 shared benchmarks; MiMo V2.6 Pro leads on 15.

The largest observed score gap is 22.45 pts on IOI , where GPT-5.6 Luna leads.

Reported ±1 standard-error ranges overlap on 9 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark GPT-5.6 Luna MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 51.69% ±1.06 55.20% ±1.18 3.51 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 1.25% ±0.83 10.83% ±2.45 9.58 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 36.54% ±3.35 47.12% ±3.47 10.58 pts Reported ±1 SE ranges do not overlap
EMB 67.12% ±2.94 62.86% ±3.14 4.26 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 55.04% ±0.31 57.34% ±0.57 2.30 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 60.81% ±3.26 64.94% ±3.21 4.12 pts Reported ±1 SE ranges overlap
MedCode 42.39% ±2.27 44.97% ±2.10 2.58 pts Reported ±1 SE ranges overlap
MedScribe 84.39% ±2.58 88.31% ±1.94 3.92 pts Reported ±1 SE ranges overlap
ProofBench v1.1 60.00% ±4.92 70.00% ±4.61 10.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 14.41% ±2.36 15.31% ±2.42 0.90 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 0.00% ±0.00 2.86% ±2.01 2.86 pts Reported ±1 SE ranges do not overlap
SAGE 44.22% ±3.33 45.05% ±3.40 0.83 pts Reported ±1 SE ranges overlap
Code Migration 44.55% ±4.24 43.01% ±4.32 1.54 pts Reported ±1 SE ranges overlap
IOI 61.78% ±11.61 39.33% ±2.41 22.45 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 11.62% ±1.01 31.31% ±3.07 19.70 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 77.06% ±3.07 85.22% ±3.39 8.17 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 73.63% ±5.56 72.86% ±5.50 0.77 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 61.16% ±1.27 68.94% ±1.20 7.78 pts Reported ±1 SE ranges do not overlap

Performance by category

Category GPT-5.6 Luna average MiMo V2.6 Pro average
Legal 18.89% 28.97%
Finance 60.99% 61.71%
Healthcare 63.39% 66.64%
Math 60.00% 70.00%
Science 7.21% 9.09%
Education 44.22% 45.05%
Coding 39.00% 39.88%
Cyber 73.63% 72.86%
Social Mobility 61.16% 68.94%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark GPT-5.6 Luna cost MiMo V2.6 Pro cost GPT-5.6 Luna latency MiMo V2.6 Pro latency
Vals Index $0.82 $0.41 29m27s 1h07m
Harvey's Legal Agent Benchmark $0.38 $0.22 10m56s 25m49s
Legal Research Bench $0.85 $0.18 39m34s 30m16s
EMB $0.37 $0.33 11m00s 55m14s
Finance Agent (v2) $0.28 $0.20 12m52s 10m48s
Tax Agent Bench $0.47 $0.13 42m46s 25m48s
MedCode N/A N/A 81.28s 2m57s
MedScribe N/A N/A 116.89s 2m53s
ProofBench v1.1 $0.12 $0.25 8m34s 1h02m
MysteryMechanism $0.11 $0.15 8m13s 44m00s
Terminal-Bench Science $0.58 $0.62 2h03m 3h12m
SAGE N/A N/A 51.88s 3m40s
Code Migration $1.88 $0.69 58m35s 2h49m
IOI $0.58 $0.84 1h10m 2h26m
ProgramBench $0.94 $0.73 41m34s 3h00m
Terminal-Bench 4.0 $0.73 $0.50 35m55s 2h42m
Vibe Code Bench v1.1 $0.73 $1.04 25m52s 1h01m
CyberBench v1.1 $0.40 $0.09 16m15s 27m41s
Public Benefits Bench v1.1 $0.27 $0.08 38m40s 37m14s

Results available only for GPT-5.6 Luna

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • BioMysteryBench
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for MiMo V2.6 Pro

  • SRE Bench
Model details GPT-5.6 Luna Model details MiMo V2.6 Pro