GPT-5.6 Terra vs MiMo V2.6 Flash: Benchmark Comparison

GPT-5.6 Terra has the higher score on 9 of 18 shared benchmarks; MiMo V2.6 Flash leads on 8.

The largest observed score gap is 39.94 pts on IOI , where GPT-5.6 Terra leads.

Reported ±1 standard-error ranges overlap on 14 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark GPT-5.6 Terra MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 53.09% ±1.29 53.23% ±1.10 0.15 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 0.83% ±0.83 11.25% ±2.54 10.42 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.35% ±3.42 37.98% ±3.37 3.36 pts Reported ±1 SE ranges overlap
EMB 66.20% ±3.11 65.46% ±2.71 0.75 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 54.44% ±2.07 56.28% ±0.46 1.84 pts Reported ±1 SE ranges overlap
Tax Agent Bench 65.20% ±3.19 59.90% ±3.13 5.29 pts Reported ±1 SE ranges overlap
MedCode 43.41% ±2.17 41.06% ±2.03 2.36 pts Reported ±1 SE ranges overlap
MedScribe 82.87% ±1.95 85.28% ±1.99 2.41 pts Reported ±1 SE ranges overlap
ProofBench v1.1 74.00% ±4.41 63.00% ±4.85 11.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 10.00% ±3.61 4.29% ±2.44 5.71 pts Reported ±1 SE ranges overlap
SAGE 47.00% ±3.40 43.53% ±3.43 3.47 pts Reported ±1 SE ranges overlap
Code Migration 47.80% ±4.28 40.93% ±4.35 6.88 pts Reported ±1 SE ranges overlap
IOI 87.61% ±6.53 47.67% ±4.79 39.94 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 22.73% ±2.31 24.24% ±1.51 1.52 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 74.59% ±4.20 78.96% ±4.04 4.37 pts Reported ±1 SE ranges overlap
CyberBench v1.1 72.08% ±5.41 75.36% ±5.42 3.27 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 62.38% ±1.26 67.59% ±1.22 5.21 pts Reported ±1 SE ranges do not overlap

Performance by category

Category GPT-5.6 Terra average MiMo V2.6 Flash average
Legal 21.09% 24.62%
Finance 61.95% 60.55%
Healthcare 63.14% 63.17%
Math 74.00% 63.00%
Science 10.00% 4.29%
Education 47.00% 43.53%
Coding 46.65% 38.46%
Cyber 72.08% 75.36%
Social Mobility 62.38% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark GPT-5.6 Terra cost MiMo V2.6 Flash cost GPT-5.6 Terra latency MiMo V2.6 Flash latency
Vals Index $5.78 $0.20 35m39s 58m02s
Harvey's Legal Agent Benchmark $3.20 $0.09 17m05s 17m12s
Legal Research Bench $7.91 $0.07 1h10m 15m20s
EMB $2.19 $0.11 13m33s 25m22s
Finance Agent (v2) $3.63 $0.07 24m10s 7m02s
Tax Agent Bench $4.83 $0.05 1h01m 17m19s
MedCode N/A N/A 18.41s 93.37s
MedScribe N/A N/A 35.53s 74.48s
ProofBench v1.1 $1.27 $0.12 12m57s 1h03m
Terminal-Bench Science $5.19 $0.24 2h04m 3h59m
SAGE N/A N/A 40.06s 2m16s
Code Migration $8.13 $0.49 48m34s 3h12m
IOI $8.66 $0.22 1h23m 1h49m
ProgramBench $5.38 $0.47 33m36s 3h22m
Terminal-Bench 4.0 $5.60 $0.22 31m29s 2h25m
Vibe Code Bench v1.1 $7.89 $0.56 19m12s 50m13s
CyberBench v1.1 $3.31 $0.05 17m04s 28m22s
Public Benefits Bench v1.1 $1.20 $0.03 19m31s 20m56s

Results available only for GPT-5.6 Terra

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for MiMo V2.6 Flash

  • BioMysteryBench
  • MysteryMechanism
  • SRE Bench
Model details GPT-5.6 Terra Model details MiMo V2.6 Flash