GPT-6 Luna vs MiMo V2.6 Pro: Benchmark Comparison

GPT-6 Luna has the higher score on 6 of 19 shared benchmarks; MiMo V2.6 Pro leads on 12.

The largest observed score gap is 17.68 pts on Terminal-Bench 4.0 , where MiMo V2.6 Pro leads.

Reported ±1 standard-error ranges overlap on 11 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark GPT-6 Luna MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 51.22% ±1.07 55.20% ±1.18 3.98 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 2.92% ±1.36 10.83% ±2.45 7.92 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 30.29% ±3.19 47.12% ±3.47 16.83 pts Reported ±1 SE ranges do not overlap
EMB 68.52% ±2.81 62.86% ±3.14 5.66 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 49.87% ±0.23 57.34% ±0.57 7.47 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 58.86% ±3.21 64.94% ±3.21 6.07 pts Reported ±1 SE ranges overlap
MedCode 44.69% ±2.30 44.97% ±2.10 0.28 pts Reported ±1 SE ranges overlap
MedScribe 83.71% ±1.95 88.31% ±1.94 4.60 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 64.00% ±4.82 70.00% ±4.61 6.00 pts Reported ±1 SE ranges overlap
MysteryMechanism 19.37% ±2.66 15.31% ±2.42 4.05 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 4.29% ±2.44 2.86% ±2.01 1.43 pts Reported ±1 SE ranges overlap
SAGE 48.09% ±3.38 45.05% ±3.40 3.04 pts Reported ±1 SE ranges overlap
Code Migration 42.55% ±4.42 43.01% ±4.32 0.46 pts Reported ±1 SE ranges overlap
IOI 55.56% ±8.87 39.33% ±2.41 16.22 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 13.64% ±1.51 31.31% ±3.07 17.68 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 81.65% ±3.38 85.22% ±3.39 3.57 pts Reported ±1 SE ranges overlap
CyberBench v1.1 76.25% ±5.29 72.86% ±5.50 3.39 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 57.65% ±1.28 68.94% ±1.20 11.30 pts Reported ±1 SE ranges do not overlap

Performance by category

Category GPT-6 Luna average MiMo V2.6 Pro average
Legal 16.60% 28.97%
Finance 59.08% 61.71%
Healthcare 64.20% 66.64%
Math 64.00% 70.00%
Science 11.83% 9.09%
Education 48.09% 45.05%
Coding 38.78% 39.88%
Cyber 76.25% 72.86%
Social Mobility 57.65% 68.94%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark GPT-6 Luna cost MiMo V2.6 Pro cost GPT-6 Luna latency MiMo V2.6 Pro latency
Vals Index $0.43 $0.41 30m34s 1h07m
Harvey's Legal Agent Benchmark $0.30 $0.22 17m32s 25m49s
Legal Research Bench $0.44 $0.18 42m28s 30m16s
EMB $0.14 $0.33 17m21s 55m14s
Finance Agent (v2) $0.12 $0.20 14m58s 10m48s
Tax Agent Bench $0.19 $0.13 33m16s 25m48s
MedCode N/A N/A 108.77s 2m57s
MedScribe N/A N/A 2m50s 2m53s
ProofBench v1.1 $0.04 $0.25 7m31s 1h02m
MysteryMechanism $0.03 $0.15 6m01s 44m00s
Terminal-Bench Science $0.24 $0.62 1h16m 3h12m
SAGE N/A N/A 91.07s 3m40s
Code Migration $0.60 $0.69 50m23s 2h49m
IOI $0.20 $0.84 35m47s 2h26m
ProgramBench $0.18 $0.73 28m55s 3h00m
Terminal-Bench 4.0 $0.35 $0.50 32m46s 2h42m
Vibe Code Bench v1.1 $1.35 $1.04 37m00s 1h01m
CyberBench v1.1 $0.13 $0.09 15m50s 27m41s
Public Benefits Bench v1.1 $0.22 $0.08 52m59s 37m14s

Results available only for GPT-6 Luna

  • BioMysteryBench

Results available only for MiMo V2.6 Pro

  • SRE Bench
Model details GPT-6 Luna Model details MiMo V2.6 Pro