DeepSeek V4.1 Flash vs MiMo V2.6 Pro: Benchmark Comparison

DeepSeek V4.1 Flash has the higher score on 5 of 20 shared benchmarks; MiMo V2.6 Pro leads on 14.

The largest observed score gap is 16.00 pts on ProofBench v1.1 , where MiMo V2.6 Pro leads.

Reported ±1 standard-error ranges overlap on 12 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4.1 Flash MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 51.32% ±1.13 55.20% ±1.18 3.88 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 6.67% ±1.83 10.83% ±2.45 4.17 pts Reported ±1 SE ranges overlap
Legal Research Bench 41.35% ±3.42 47.12% ±3.47 5.77 pts Reported ±1 SE ranges overlap
EMB 57.21% ±3.23 62.86% ±3.14 5.65 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.48% ±0.39 57.34% ±0.57 3.86 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.46% ±3.15 64.94% ±3.21 2.48 pts Reported ±1 SE ranges overlap
MedCode 41.17% ±2.04 44.97% ±2.10 3.80 pts Reported ±1 SE ranges overlap
MedScribe 85.50% ±1.92 88.31% ±1.94 2.81 pts Reported ±1 SE ranges overlap
ProofBench v1.1 54.00% ±5.01 70.00% ±4.61 16.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 21.17% ±2.75 15.31% ±2.42 5.86 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 0.00% ±0.00 2.86% ±2.01 2.86 pts Reported ±1 SE ranges do not overlap
SAGE 47.88% ±3.44 45.05% ±3.40 2.83 pts Reported ±1 SE ranges overlap
Code Migration 45.62% ±4.29 43.01% ±4.32 2.61 pts Reported ±1 SE ranges overlap
IOI 40.28% ±2.56 39.33% ±2.41 0.95 pts Reported ±1 SE ranges overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 19.70% ±1.75 31.31% ±3.07 11.62 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 84.74% ±2.89 85.22% ±3.39 0.48 pts Reported ±1 SE ranges overlap
CyberBench v1.1 73.69% ±5.48 72.86% ±5.50 0.83 pts Reported ±1 SE ranges overlap
SRE Bench 0.76% ±0.54 3.05% ±1.06 2.29 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 64.28% ±1.25 68.94% ±1.20 4.67 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4.1 Flash average MiMo V2.6 Pro average
Legal 24.01% 28.97%
Finance 57.72% 61.71%
Healthcare 63.34% 66.64%
Math 54.00% 70.00%
Science 10.59% 9.09%
Education 47.88% 45.05%
Coding 38.17% 39.88%
Cyber 37.23% 37.95%
Social Mobility 64.28% 68.94%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4.1 Flash cost MiMo V2.6 Pro cost DeepSeek V4.1 Flash latency MiMo V2.6 Pro latency
Vals Index $0.33 $0.41 27m53s 1h07m
Harvey's Legal Agent Benchmark $0.17 $0.22 9m39s 25m49s
Legal Research Bench $0.25 $0.18 11m42s 30m16s
EMB $0.23 $0.33 13m00s 55m14s
Finance Agent (v2) $0.21 $0.20 6m04s 10m48s
Tax Agent Bench $0.13 $0.13 5m49s 25m48s
MedCode N/A N/A 35.13s 2m57s
MedScribe N/A N/A 53.83s 2m53s
ProofBench v1.1 $0.13 $0.25 9m07s 1h02m
MysteryMechanism $0.16 $0.15 8m04s 44m00s
Terminal-Bench Science $0.47 $0.62 2h57m 3h12m
SAGE N/A N/A 35.78s 3m40s
Code Migration $0.94 $0.69 1h45m 2h49m
IOI $0.24 $0.84 18m30s 2h26m
ProgramBench $0.90 $0.73 1h02m 3h00m
Terminal-Bench 4.0 $0.50 $0.50 48m22s 2h42m
Vibe Code Bench v1.1 $0.41 $1.04 15m28s 1h01m
CyberBench v1.1 $0.07 $0.09 13m50s 27m41s
SRE Bench $0.55 $0.53 1h01m 2h23m
Public Benefits Bench v1.1 $0.07 $0.08 9m59s 37m14s

Results available only for DeepSeek V4.1 Flash

  • LegalBench
  • BioMysteryBench
  • SkillsBench
  • Vibe Code Bench 1-100

Results available only for MiMo V2.6 Pro

None.

Model details DeepSeek V4.1 Flash Model details MiMo V2.6 Pro