GPT-5.6 Terra vs MiMo V2.6 Pro: Benchmark Comparison

GPT-5.6 Terra has the higher score on 7 of 18 shared benchmarks; MiMo V2.6 Pro leads on 10.

The largest observed score gap is 48.28 pts on IOI , where GPT-5.6 Terra leads.

Reported ±1 standard-error ranges overlap on 10 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark GPT-5.6 Terra MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 53.09% ±1.29 55.20% ±1.18 2.11 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 0.83% ±0.83 10.83% ±2.45 10.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.35% ±3.42 47.12% ±3.47 5.77 pts Reported ±1 SE ranges overlap
EMB 66.20% ±3.11 62.86% ±3.14 3.35 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 54.44% ±2.07 57.34% ±0.57 2.90 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 65.20% ±3.19 64.94% ±3.21 0.26 pts Reported ±1 SE ranges overlap
MedCode 43.41% ±2.17 44.97% ±2.10 1.55 pts Reported ±1 SE ranges overlap
MedScribe 82.87% ±1.95 88.31% ±1.94 5.44 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 74.00% ±4.41 70.00% ±4.61 4.00 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 10.00% ±3.61 2.86% ±2.01 7.14 pts Reported ±1 SE ranges do not overlap
SAGE 47.00% ±3.40 45.05% ±3.40 1.95 pts Reported ±1 SE ranges overlap
Code Migration 47.80% ±4.28 43.01% ±4.32 4.79 pts Reported ±1 SE ranges overlap
IOI 87.61% ±6.53 39.33% ±2.41 48.28 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 22.73% ±2.31 31.31% ±3.07 8.59 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 74.59% ±4.20 85.22% ±3.39 10.63 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 72.08% ±5.41 72.86% ±5.50 0.77 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 62.38% ±1.26 68.94% ±1.20 6.56 pts Reported ±1 SE ranges do not overlap

Performance by category

Category GPT-5.6 Terra average MiMo V2.6 Pro average
Legal 21.09% 28.97%
Finance 61.95% 61.71%
Healthcare 63.14% 66.64%
Math 74.00% 70.00%
Science 10.00% 2.86%
Education 47.00% 45.05%
Coding 46.65% 39.88%
Cyber 72.08% 72.86%
Social Mobility 62.38% 68.94%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark GPT-5.6 Terra cost MiMo V2.6 Pro cost GPT-5.6 Terra latency MiMo V2.6 Pro latency
Vals Index $5.78 $0.41 35m39s 1h07m
Harvey's Legal Agent Benchmark $3.20 $0.22 17m05s 25m49s
Legal Research Bench $7.91 $0.18 1h10m 30m16s
EMB $2.19 $0.33 13m33s 55m14s
Finance Agent (v2) $3.63 $0.20 24m10s 10m48s
Tax Agent Bench $4.83 $0.13 1h01m 25m48s
MedCode N/A N/A 18.41s 2m57s
MedScribe N/A N/A 35.53s 2m53s
ProofBench v1.1 $1.27 $0.25 12m57s 1h02m
Terminal-Bench Science $5.19 $0.62 2h04m 3h12m
SAGE N/A N/A 40.06s 3m40s
Code Migration $8.13 $0.69 48m34s 2h49m
IOI $8.66 $0.84 1h23m 2h26m
ProgramBench $5.38 $0.73 33m36s 3h00m
Terminal-Bench 4.0 $5.60 $0.50 31m29s 2h42m
Vibe Code Bench v1.1 $7.89 $1.04 19m12s 1h01m
CyberBench v1.1 $3.31 $0.09 17m04s 27m41s
Public Benefits Bench v1.1 $1.20 $0.08 19m31s 37m14s

Results available only for GPT-5.6 Terra

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for MiMo V2.6 Pro

  • MysteryMechanism
  • SRE Bench
Model details GPT-5.6 Terra Model details MiMo V2.6 Pro