GPT-5.6 Luna vs MiMo V2.6 Flash: Benchmark Comparison

GPT-5.6 Luna has the higher score on 6 of 20 shared benchmarks; MiMo V2.6 Flash leads on 14.

The largest observed score gap is 14.11 pts on IOI , where GPT-5.6 Luna leads.

Reported ±1 standard-error ranges overlap on 13 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark GPT-5.6 Luna MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 51.69% ±1.06 53.23% ±1.10 1.54 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 1.25% ±0.83 11.25% ±2.54 10.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 36.54% ±3.35 37.98% ±3.37 1.44 pts Reported ±1 SE ranges overlap
EMB 67.12% ±2.94 65.46% ±2.71 1.66 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 55.04% ±0.31 56.28% ±0.46 1.23 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 60.81% ±3.26 59.90% ±3.13 0.91 pts Reported ±1 SE ranges overlap
MedCode 42.39% ±2.27 41.06% ±2.03 1.33 pts Reported ±1 SE ranges overlap
MedScribe 84.39% ±2.58 85.28% ±1.99 0.88 pts Reported ±1 SE ranges overlap
ProofBench v1.1 60.00% ±4.92 63.00% ±4.85 3.00 pts Reported ±1 SE ranges overlap
BioMysteryBench 61.48% ±0.74 69.26% ±0.98 7.78 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 14.41% ±2.36 21.62% ±2.77 7.21 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 0.00% ±0.00 4.29% ±2.44 4.29 pts Reported ±1 SE ranges do not overlap
SAGE 44.22% ±3.33 43.53% ±3.43 0.69 pts Reported ±1 SE ranges overlap
Code Migration 44.55% ±4.24 40.93% ±4.35 3.62 pts Reported ±1 SE ranges overlap
IOI 61.78% ±11.61 47.67% ±4.79 14.11 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 11.62% ±1.01 24.24% ±1.51 12.63 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 77.06% ±3.07 78.96% ±4.04 1.90 pts Reported ±1 SE ranges overlap
CyberBench v1.1 73.63% ±5.56 75.36% ±5.42 1.73 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 61.16% ±1.27 67.59% ±1.22 6.43 pts Reported ±1 SE ranges do not overlap

Performance by category

Category GPT-5.6 Luna average MiMo V2.6 Flash average
Legal 18.89% 24.62%
Finance 60.99% 60.55%
Healthcare 63.39% 63.17%
Math 60.00% 63.00%
Science 25.30% 31.72%
Education 44.22% 43.53%
Coding 39.00% 38.46%
Cyber 73.63% 75.36%
Social Mobility 61.16% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark GPT-5.6 Luna cost MiMo V2.6 Flash cost GPT-5.6 Luna latency MiMo V2.6 Flash latency
Vals Index $0.82 $0.20 29m27s 58m02s
Harvey's Legal Agent Benchmark $0.38 $0.09 10m56s 17m12s
Legal Research Bench $0.85 $0.07 39m34s 15m20s
EMB $0.37 $0.11 11m00s 25m22s
Finance Agent (v2) $0.28 $0.07 12m52s 7m02s
Tax Agent Bench $0.47 $0.05 42m46s 17m19s
MedCode N/A N/A 81.28s 93.37s
MedScribe N/A N/A 116.89s 74.48s
ProofBench v1.1 $0.12 $0.12 8m34s 1h03m
BioMysteryBench $0.09 $0.05 18m31s 28m28s
MysteryMechanism $0.11 $0.06 8m13s 35m08s
Terminal-Bench Science $0.58 $0.24 2h03m 3h59m
SAGE N/A N/A 51.88s 2m16s
Code Migration $1.88 $0.49 58m35s 3h12m
IOI $0.58 $0.22 1h10m 1h49m
ProgramBench $0.94 $0.47 41m34s 3h22m
Terminal-Bench 4.0 $0.73 $0.22 35m55s 2h25m
Vibe Code Bench v1.1 $0.73 $0.56 25m52s 50m13s
CyberBench v1.1 $0.40 $0.05 16m15s 28m22s
Public Benefits Bench v1.1 $0.27 $0.03 38m40s 20m56s

Results available only for GPT-5.6 Luna

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for MiMo V2.6 Flash

  • SRE Bench
Model details GPT-5.6 Luna Model details MiMo V2.6 Flash