MiniMax-M3 vs GPT 5.4 Nano: Benchmark Comparison
MiniMax-M3 has the higher score on 17 of 18 shared benchmarks; GPT 5.4 Nano leads on 1.
The largest observed score gap is 23.56 pts on Legal Research Bench , where MiniMax-M3 leads.
Reported ±1 standard-error ranges overlap on 3 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | MiniMax-M3 | GPT 5.4 Nano | Gap | Reported uncertainty |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | 4.17% ±1.65 | 0.00% ±0.00 | 4.17 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 29.81% ±3.18 | 6.25% ±1.68 | 23.56 pts | Reported ±1 SE ranges do not overlap |
| LegalBench | 85.42% ±0.44 | 77.92% ±0.43 | 7.50 pts | Reported ±1 SE ranges do not overlap |
| EMB | 47.76% ±2.95 | 44.75% ±3.12 | 3.01 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 48.27% ±0.44 | 38.22% ±1.19 | 10.05 pts | Reported ±1 SE ranges do not overlap |
| MortgageTax | 68.36% ±0.91 | 59.10% ±0.96 | 9.26 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 49.69% ±3.25 | 26.58% ±2.52 | 23.11 pts | Reported ±1 SE ranges do not overlap |
| TaxEval v2 | 72.73% ±0.86 | 67.42% ±0.93 | 5.31 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 46.29% ±2.10 | 41.03% ±2.26 | 5.26 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 87.25% ±1.96 | 77.09% ±1.89 | 10.16 pts | Reported ±1 SE ranges do not overlap |
| GPQA Diamond | 92.68% ±1.44 | 77.53% ±2.42 | 15.15 pts | Reported ±1 SE ranges do not overlap |
| MMLU Pro | 84.22% ±0.36 | 77.17% ±0.43 | 7.05 pts | Reported ±1 SE ranges do not overlap |
| MMMU Pro | 81.16% ±0.94 | 73.58% ±1.06 | 7.57 pts | Reported ±1 SE ranges do not overlap |
| SAGE | 50.57% ±3.44 | 38.08% ±3.09 | 12.49 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 19.93% ±3.94 | 14.47% ±4.03 | 5.46 pts | Reported ±1 SE ranges overlap |
| LiveCodeBench | 82.15% ±1.05 | 84.01% ±1.04 | 1.86 pts | Reported ±1 SE ranges overlap |
| SWE-bench | 75.00% ±1.94 | 69.80% ±2.06 | 5.20 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 47.57% ±5.44 | 26.10% ±5.08 | 21.47 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | MiniMax-M3 average | GPT 5.4 Nano average |
|---|---|---|
| Legal | 39.80% | 28.06% |
| Finance | 57.36% | 47.21% |
| Healthcare | 66.77% | 59.06% |
| Academic | 86.02% | 76.09% |
| Education | 50.57% | 38.08% |
| Coding | 56.16% | 48.59% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | MiniMax-M3 cost | GPT 5.4 Nano cost | MiniMax-M3 latency | GPT 5.4 Nano latency |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | $1.46 | $0.18 | 22m31s | 6m03s |
| Legal Research Bench | $0.34 | $0.14 | 13m34s | 8m47s |
| LegalBench | N/A | N/A | 8.21s | 2.70s |
| EMB | $2.09 | $0.57 | 31m30s | 29m27s |
| Finance Agent (v2) | $0.32 | $0.16 | 8m17s | 5m35s |
| MortgageTax | N/A | N/A | 26.76s | 14.53s |
| Tax Agent Bench | $0.16 | $0.07 | 5m15s | 3m40s |
| TaxEval v2 | N/A | N/A | 94.04s | 15.90s |
| MedCode | N/A | N/A | 63.12s | 8.04s |
| MedScribe | N/A | N/A | 2m04s | 20.53s |
| GPQA Diamond | N/A | N/A | 4m39s | 29.50s |
| MMLU Pro | N/A | N/A | 41.39s | 8.21s |
| MMMU Pro | N/A | N/A | 68.35s | 21.91s |
| SAGE | N/A | N/A | 118.54s | 34.33s |
| Code Migration | $7.07 | $0.42 | 1h14m | 26m34s |
| LiveCodeBench | N/A | N/A | 6m07s | 60.09s |
| SWE-bench | $0.42 | $0.10 | 12m06s | 4m23s |
| Vibe Code Bench v1.1 | $6.45 | $1.28 | 1h25m | 54m36s |
Results available only for MiniMax-M3
- Vals Index
- ProofBench v1.1
- SkillsBench
- Terminal-Bench 4.0
- Vibe Code Bench 1-100
- CyberBench v1.1
- Public Benefits Bench v1.1
Results available only for GPT 5.4 Nano
None.