Inkling vs MiMo V2.5: Benchmark Comparison

Inkling has the higher score on 13 of 19 shared benchmarks; MiMo V2.5 leads on 6.

The largest observed score gap is 22.96 pts on Vibe Code Bench v1.1 , where MiMo V2.5 leads.

Reported ±1 standard-error ranges overlap on 3 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Inkling MiMo V2.5 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 2.08% ±0.83 1.67% ±0.00 0.42 pts Reported ±1 SE ranges overlap
Legal Research Bench 28.36% ±3.13 9.13% ±2.00 19.23 pts Reported ±1 SE ranges do not overlap
LegalBench 82.95% ±0.50 78.89% ±0.51 4.06 pts Reported ±1 SE ranges do not overlap
EMB 38.36% ±3.27 55.09% ±3.13 16.73 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 46.60% ±0.93 36.73% ±0.11 9.87 pts Reported ±1 SE ranges do not overlap
MortgageTax 63.99% ±0.94 59.26% ±0.98 4.73 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 43.52% ±2.95 29.16% ±2.71 14.35 pts Reported ±1 SE ranges do not overlap
TaxEval v2 75.31% ±0.84 71.83% ±0.88 3.47 pts Reported ±1 SE ranges do not overlap
MedCode 41.19% ±2.23 31.89% ±2.02 9.29 pts Reported ±1 SE ranges do not overlap
MedScribe 85.41% ±1.84 72.15% ±1.85 13.25 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 0.00% ±0.00 16.00% ±3.69 16.00 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 87.12% ±1.78 81.57% ±2.05 5.56 pts Reported ±1 SE ranges do not overlap
MMLU Pro 86.30% ±0.34 82.93% ±0.37 3.36 pts Reported ±1 SE ranges do not overlap
MMMU Pro 78.50% ±0.99 80.00% ±0.96 1.50 pts Reported ±1 SE ranges overlap
SAGE 36.55% ±3.22 43.27% ±3.39 6.71 pts Reported ±1 SE ranges do not overlap
Code Migration 11.79% ±3.40 14.23% ±3.69 2.44 pts Reported ±1 SE ranges overlap
LiveCodeBench 85.52% ±1.00 81.51% ±1.07 4.00 pts Reported ±1 SE ranges do not overlap
SWE-bench 77.60% ±1.87 71.00% ±2.03 6.60 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 19.21% ±3.66 42.17% ±4.57 22.96 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Inkling average MiMo V2.5 average
Legal 37.80% 29.90%
Finance 53.55% 50.41%
Healthcare 63.30% 52.02%
Math 0.00% 16.00%
Academic 83.97% 81.50%
Education 36.55% 43.27%
Coding 48.53% 52.23%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Inkling cost MiMo V2.5 cost Inkling latency MiMo V2.5 latency
Harvey's Legal Agent Benchmark $1.25 $0.04 18m24s 6m59s
Legal Research Bench $0.49 $0.03 14m29s 6m31s
LegalBench N/A N/A 23.14s 11.18s
EMB $1.67 $0.08 26m22s 30m38s
Finance Agent (v2) $1.00 $0.09 18m02s 5m39s
MortgageTax N/A N/A 69.11s 32.94s
Tax Agent Bench $0.35 $0.02 13m31s 8m21s
TaxEval v2 N/A N/A 3m18s 26.30s
MedCode N/A N/A 2m47s 17.70s
MedScribe N/A N/A 4m05s 20.26s
ProofBench v1.1 $0.20 $0.06 7m16s 20m56s
GPQA Diamond N/A N/A 5m09s 89.39s
MMLU Pro N/A N/A 73.26s 30.25s
MMMU Pro N/A N/A 2m04s 42.47s
SAGE N/A N/A 3m18s 52.66s
Code Migration $3.01 $0.10 29m34s 37m55s
LiveCodeBench N/A N/A 4m18s 108.61s
SWE-bench $0.59 $0.01 6m34s 4m04s
Vibe Code Bench v1.1 $1.61 $0.07 18m25s 26m20s

Results available only for Inkling

  • Vals Index
  • IOI
  • ProgramBench
  • SkillsBench
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • Public Benefits Bench v1.1

Results available only for MiMo V2.5

None.

Model details Inkling Model details MiMo V2.5