DeepSeek V4 vs MiniMax-M3: Benchmark Comparison

DeepSeek V4 has the higher score on 9 of 20 shared benchmarks; MiniMax-M3 leads on 11.

The largest observed score gap is 12.11 pts on MedScribe , where MiniMax-M3 leads.

Reported ±1 standard-error ranges overlap on 10 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4 MiniMax-M3 Gap Reported uncertainty
Vals Index 38.63% ±1.16 36.53% ±1.17 2.10 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 3.75% ±1.43 4.17% ±1.65 0.42 pts Reported ±1 SE ranges overlap
Legal Research Bench 23.08% ±2.93 29.81% ±3.18 6.73 pts Reported ±1 SE ranges do not overlap
LegalBench 80.32% ±0.47 85.42% ±0.44 5.09 pts Reported ±1 SE ranges do not overlap
EMB 51.62% ±3.05 47.76% ±2.95 3.86 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 44.08% ±0.65 48.27% ±0.44 4.19 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 58.50% ±3.13 49.69% ±3.25 8.81 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.08% ±0.88 72.73% ±0.86 0.65 pts Reported ±1 SE ranges overlap
MedCode 40.45% ±2.12 46.29% ±2.10 5.83 pts Reported ±1 SE ranges do not overlap
MedScribe 75.14% ±2.00 87.25% ±1.96 12.11 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 16.00% ±3.69 18.00% ±3.86 2.00 pts Reported ±1 SE ranges overlap
GPQA Diamond 89.39% ±1.65 92.68% ±1.44 3.28 pts Reported ±1 SE ranges do not overlap
MMLU Pro 87.25% ±0.34 84.22% ±0.36 3.03 pts Reported ±1 SE ranges do not overlap
Code Migration 26.20% ±4.04 19.93% ±3.94 6.28 pts Reported ±1 SE ranges overlap
LiveCodeBench 87.48% ±0.95 82.15% ±1.05 5.33 pts Reported ±1 SE ranges do not overlap
SkillsBench 51.27% ±4.60 51.50% ±4.50 0.22 pts Reported ±1 SE ranges overlap
SWE-bench 77.40% ±1.87 75.00% ±1.94 2.40 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 11.11% ±1.34 1.01% ±1.01 10.10 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 49.93% ±4.77 47.57% ±5.44 2.36 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 62.92% ±1.26 64.14% ±1.25 1.22 pts Reported ±1 SE ranges overlap

Performance by category

Category DeepSeek V4 average MiniMax-M3 average
Legal 35.72% 39.80%
Finance 56.57% 54.61%
Healthcare 57.80% 66.77%
Math 16.00% 18.00%
Academic 88.32% 88.45%
Coding 50.57% 46.19%
Social Mobility 62.92% 64.14%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4 cost MiniMax-M3 cost DeepSeek V4 latency MiniMax-M3 latency
Vals Index $1.31 $3.08 32m09s 38m55s
Harvey's Legal Agent Benchmark $0.74 $1.46 6m42s 22m31s
Legal Research Bench $0.68 $0.34 17m05s 13m34s
LegalBench N/A N/A 3m04s 8.21s
EMB $0.80 $2.09 13m42s 31m30s
Finance Agent (v2) $0.88 $0.32 25m26s 8m17s
Tax Agent Bench $0.72 $0.16 20m45s 5m15s
TaxEval v2 N/A N/A 3m10s 94.04s
MedCode N/A N/A 6m23s 63.12s
MedScribe N/A N/A 5m46s 2m04s
ProofBench v1.1 $0.03 $0.42 7m03s 11m59s
GPQA Diamond N/A N/A 15m00s 4m39s
MMLU Pro N/A N/A 9m38s 41.39s
Code Migration $1.07 $7.07 28m58s 1h14m
LiveCodeBench N/A N/A 11m56s 6m07s
SkillsBench $0.40 $0.50 8m16s 15m28s
SWE-bench $0.44 $0.42 10m35s 12m06s
Terminal-Bench 4.0 $3.38 $7.23 1h28m 1h16m
Vibe Code Bench v1.1 $2.21 $6.45 56m59s 1h25m
Public Benefits Bench v1.1 $0.46 $0.25 15m22s 11m31s

Results available only for DeepSeek V4

  • ProgramBench

Results available only for MiniMax-M3

  • MortgageTax
  • MMMU Pro
  • SAGE
  • Vibe Code Bench 1-100
  • CyberBench v1.1
Model details DeepSeek V4 Model details MiniMax-M3