GPT-6 Luna vs MiMo V2.6 Flash: Benchmark Comparison

GPT-6 Luna has the higher score on 8 of 20 shared benchmarks; MiMo V2.6 Flash leads on 10.

The largest observed score gap is 10.61 pts on Terminal-Bench 4.0 , where MiMo V2.6 Flash leads.

Reported ±1 standard-error ranges overlap on 14 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark GPT-6 Luna MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 51.22% ±1.07 53.23% ±1.10 2.02 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 2.92% ±1.36 11.25% ±2.54 8.33 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 30.29% ±3.19 37.98% ±3.37 7.69 pts Reported ±1 SE ranges do not overlap
EMB 68.52% ±2.81 65.46% ±2.71 3.06 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 49.87% ±0.23 56.28% ±0.46 6.40 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 58.86% ±3.21 59.90% ±3.13 1.04 pts Reported ±1 SE ranges overlap
MedCode 44.69% ±2.30 41.06% ±2.03 3.63 pts Reported ±1 SE ranges overlap
MedScribe 83.71% ±1.95 85.28% ±1.99 1.57 pts Reported ±1 SE ranges overlap
ProofBench v1.1 64.00% ±4.82 63.00% ±4.85 1.00 pts Reported ±1 SE ranges overlap
BioMysteryBench 61.48% ±2.59 69.26% ±0.98 7.78 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 19.37% ±2.66 21.62% ±2.77 2.25 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 4.29% ±2.44 4.29% ±2.44 0.00 pts Reported ±1 SE ranges overlap
SAGE 48.09% ±3.38 43.53% ±3.43 4.56 pts Reported ±1 SE ranges overlap
Code Migration 42.55% ±4.42 40.93% ±4.35 1.63 pts Reported ±1 SE ranges overlap
IOI 55.56% ±8.87 47.67% ±4.79 7.89 pts Reported ±1 SE ranges overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 13.64% ±1.51 24.24% ±1.51 10.61 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 81.65% ±3.38 78.96% ±4.04 2.69 pts Reported ±1 SE ranges overlap
CyberBench v1.1 76.25% ±5.29 75.36% ±5.42 0.89 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 57.65% ±1.28 67.59% ±1.22 9.95 pts Reported ±1 SE ranges do not overlap

Performance by category

Category GPT-6 Luna average MiMo V2.6 Flash average
Legal 16.60% 24.62%
Finance 59.08% 60.55%
Healthcare 64.20% 63.17%
Math 64.00% 63.00%
Science 28.38% 31.72%
Education 48.09% 43.53%
Coding 38.78% 38.46%
Cyber 76.25% 75.36%
Social Mobility 57.65% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark GPT-6 Luna cost MiMo V2.6 Flash cost GPT-6 Luna latency MiMo V2.6 Flash latency
Vals Index $0.43 $0.20 30m34s 58m02s
Harvey's Legal Agent Benchmark $0.30 $0.09 17m32s 17m12s
Legal Research Bench $0.44 $0.07 42m28s 15m20s
EMB $0.14 $0.11 17m21s 25m22s
Finance Agent (v2) $0.12 $0.07 14m58s 7m02s
Tax Agent Bench $0.19 $0.05 33m16s 17m19s
MedCode N/A N/A 108.77s 93.37s
MedScribe N/A N/A 2m50s 74.48s
ProofBench v1.1 $0.04 $0.12 7m31s 1h03m
BioMysteryBench $0.05 $0.05 8m44s 28m28s
MysteryMechanism $0.03 $0.06 6m01s 35m08s
Terminal-Bench Science $0.24 $0.24 1h16m 3h59m
SAGE N/A N/A 91.07s 2m16s
Code Migration $0.60 $0.49 50m23s 3h12m
IOI $0.20 $0.22 35m47s 1h49m
ProgramBench $0.18 $0.47 28m55s 3h22m
Terminal-Bench 4.0 $0.35 $0.22 32m46s 2h25m
Vibe Code Bench v1.1 $1.35 $0.56 37m00s 50m13s
CyberBench v1.1 $0.13 $0.05 15m50s 28m22s
Public Benefits Bench v1.1 $0.22 $0.03 52m59s 20m56s

Results available only for GPT-6 Luna

None.

Results available only for MiMo V2.6 Flash

  • SRE Bench
Model details GPT-6 Luna Model details MiMo V2.6 Flash