MiniMax-M3 vs GPT 5.4 Nano: Benchmark Comparison

MiniMax-M3 has the higher score on 17 of 18 shared benchmarks; GPT 5.4 Nano leads on 1.

The largest observed score gap is 23.56 pts on Legal Research Bench , where MiniMax-M3 leads.

Reported ±1 standard-error ranges overlap on 3 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark MiniMax-M3 GPT 5.4 Nano Gap Reported uncertainty
Harvey's Legal Agent Benchmark 4.17% ±1.65 0.00% ±0.00 4.17 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 29.81% ±3.18 6.25% ±1.68 23.56 pts Reported ±1 SE ranges do not overlap
LegalBench 85.42% ±0.44 77.92% ±0.43 7.50 pts Reported ±1 SE ranges do not overlap
EMB 47.76% ±2.95 44.75% ±3.12 3.01 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 48.27% ±0.44 38.22% ±1.19 10.05 pts Reported ±1 SE ranges do not overlap
MortgageTax 68.36% ±0.91 59.10% ±0.96 9.26 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 49.69% ±3.25 26.58% ±2.52 23.11 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.73% ±0.86 67.42% ±0.93 5.31 pts Reported ±1 SE ranges do not overlap
MedCode 46.29% ±2.10 41.03% ±2.26 5.26 pts Reported ±1 SE ranges do not overlap
MedScribe 87.25% ±1.96 77.09% ±1.89 10.16 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 92.68% ±1.44 77.53% ±2.42 15.15 pts Reported ±1 SE ranges do not overlap
MMLU Pro 84.22% ±0.36 77.17% ±0.43 7.05 pts Reported ±1 SE ranges do not overlap
MMMU Pro 81.16% ±0.94 73.58% ±1.06 7.57 pts Reported ±1 SE ranges do not overlap
SAGE 50.57% ±3.44 38.08% ±3.09 12.49 pts Reported ±1 SE ranges do not overlap
Code Migration 19.93% ±3.94 14.47% ±4.03 5.46 pts Reported ±1 SE ranges overlap
LiveCodeBench 82.15% ±1.05 84.01% ±1.04 1.86 pts Reported ±1 SE ranges overlap
SWE-bench 75.00% ±1.94 69.80% ±2.06 5.20 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 47.57% ±5.44 26.10% ±5.08 21.47 pts Reported ±1 SE ranges do not overlap

Performance by category

Category MiniMax-M3 average GPT 5.4 Nano average
Legal 39.80% 28.06%
Finance 57.36% 47.21%
Healthcare 66.77% 59.06%
Academic 86.02% 76.09%
Education 50.57% 38.08%
Coding 56.16% 48.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark MiniMax-M3 cost GPT 5.4 Nano cost MiniMax-M3 latency GPT 5.4 Nano latency
Harvey's Legal Agent Benchmark $1.46 $0.18 22m31s 6m03s
Legal Research Bench $0.34 $0.14 13m34s 8m47s
LegalBench N/A N/A 8.21s 2.70s
EMB $2.09 $0.57 31m30s 29m27s
Finance Agent (v2) $0.32 $0.16 8m17s 5m35s
MortgageTax N/A N/A 26.76s 14.53s
Tax Agent Bench $0.16 $0.07 5m15s 3m40s
TaxEval v2 N/A N/A 94.04s 15.90s
MedCode N/A N/A 63.12s 8.04s
MedScribe N/A N/A 2m04s 20.53s
GPQA Diamond N/A N/A 4m39s 29.50s
MMLU Pro N/A N/A 41.39s 8.21s
MMMU Pro N/A N/A 68.35s 21.91s
SAGE N/A N/A 118.54s 34.33s
Code Migration $7.07 $0.42 1h14m 26m34s
LiveCodeBench N/A N/A 6m07s 60.09s
SWE-bench $0.42 $0.10 12m06s 4m23s
Vibe Code Bench v1.1 $6.45 $1.28 1h25m 54m36s

Results available only for MiniMax-M3

  • Vals Index
  • ProofBench v1.1
  • SkillsBench
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • Public Benefits Bench v1.1

Results available only for GPT 5.4 Nano

None.

Model details MiniMax-M3 Model details GPT 5.4 Nano