MiniMax-M3 vs GPT 5.4 Mini: Benchmark Comparison

MiniMax-M3 has the higher score on 13 of 17 shared benchmarks; GPT 5.4 Mini leads on 4.

The largest observed score gap is 17.31 pts on Legal Research Bench , where MiniMax-M3 leads.

Reported ±1 standard-error ranges overlap on 10 of 17 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark MiniMax-M3 GPT 5.4 Mini Gap Reported uncertainty
Vals Index 36.53% ±1.17 33.17% ±1.28 3.37 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 4.17% ±1.65 0.00% ±0.00 4.17 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 29.81% ±3.18 12.50% ±2.30 17.31 pts Reported ±1 SE ranges do not overlap
EMB 47.76% ±2.95 45.43% ±3.65 2.33 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 48.27% ±0.44 45.36% ±0.45 2.91 pts Reported ±1 SE ranges do not overlap
MortgageTax 68.36% ±0.91 63.51% ±0.91 4.85 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 49.69% ±3.25 36.59% ±3.10 13.11 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.73% ±0.86 71.22% ±0.90 1.51 pts Reported ±1 SE ranges overlap
GPQA Diamond 92.68% ±1.44 83.08% ±2.46 9.60 pts Reported ±1 SE ranges do not overlap
MMLU Pro 84.22% ±0.36 84.55% ±0.36 0.33 pts Reported ±1 SE ranges overlap
MMMU Pro 81.16% ±0.94 79.25% ±0.97 1.91 pts Reported ±1 SE ranges overlap
SAGE 50.57% ±3.44 50.81% ±3.40 0.24 pts Reported ±1 SE ranges overlap
Code Migration 19.93% ±3.94 12.94% ±3.64 6.99 pts Reported ±1 SE ranges overlap
LiveCodeBench 82.15% ±1.05 81.47% ±1.09 0.69 pts Reported ±1 SE ranges overlap
SWE-bench 75.00% ±1.94 73.00% ±1.99 2.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 1.01% ±1.01 2.52% ±1.01 1.51 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 47.57% ±5.44 47.97% ±5.61 0.40 pts Reported ±1 SE ranges overlap

Performance by category

Category MiniMax-M3 average GPT 5.4 Mini average
Legal 16.99% 6.25%
Finance 57.36% 52.42%
Academic 86.02% 82.29%
Education 50.57% 50.81%
Coding 45.13% 43.58%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark MiniMax-M3 cost GPT 5.4 Mini cost MiniMax-M3 latency GPT 5.4 Mini latency
Vals Index $3.08 $1.45 38m55s 29m49s
Harvey's Legal Agent Benchmark $1.46 $1.00 22m31s 11m21s
Legal Research Bench $0.34 $1.86 13m34s 37m36s
EMB $2.09 $2.15 31m30s 43m21s
Finance Agent (v2) $0.32 $1.20 8m17s 35m54s
MortgageTax N/A N/A 26.76s 2m06s
Tax Agent Bench $0.16 $0.81 5m15s 16m46s
TaxEval v2 N/A N/A 94.04s 44.58s
GPQA Diamond N/A N/A 4m39s 63.11s
MMLU Pro N/A N/A 41.39s 20.00s
MMMU Pro N/A N/A 68.35s 50.81s
SAGE N/A N/A 118.54s 2m21s
Code Migration $7.07 $1.95 1h14m 28m53s
LiveCodeBench N/A N/A 6m07s 3m09s
SWE-bench $0.42 $0.51 12m06s 5m26s
Terminal-Bench 4.0 $7.23 $1.43 1h16m 29m41s
Vibe Code Bench v1.1 $6.45 $1.19 1h25m 34m15s

Results available only for MiniMax-M3

  • LegalBench
  • MedCode
  • MedScribe
  • ProofBench v1.1
  • SkillsBench
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • Public Benefits Bench v1.1

Results available only for GPT 5.4 Mini

  • ProgramBench
Model details MiniMax-M3 Model details GPT 5.4 Mini