Qwen 3.7 Plus vs MiniMax-M3: Benchmark Comparison
Qwen 3.7 Plus has the higher score on 2 of 10 shared benchmarks; MiniMax-M3 leads on 8.
The largest observed score gap is 13.46 pts on Legal Research Bench , where MiniMax-M3 leads.
Reported ±1 standard-error ranges overlap on 3 of 10 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Qwen 3.7 Plus | MiniMax-M3 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | 0.00% ±0.00 | 4.17% ±1.65 | 4.17 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 16.35% ±2.57 | 29.81% ±3.18 | 13.46 pts | Reported ±1 SE ranges do not overlap |
| EMB | 49.34% ±3.18 | 47.76% ±2.95 | 1.59 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 38.22% ±1.04 | 48.27% ±0.44 | 10.05 pts | Reported ±1 SE ranges do not overlap |
| MortgageTax | 66.18% ±0.93 | 68.36% ±0.91 | 2.19 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 38.71% ±2.81 | 49.69% ±3.25 | 10.98 pts | Reported ±1 SE ranges do not overlap |
| SAGE | 39.25% ±3.38 | 50.57% ±3.44 | 11.32 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 12.86% ±2.93 | 19.93% ±3.94 | 7.07 pts | Reported ±1 SE ranges do not overlap |
| SkillsBench | 54.30% ±4.33 | 51.50% ±4.50 | 2.80 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 46.39% ±4.61 | 47.57% ±5.44 | 1.18 pts | Reported ±1 SE ranges overlap |
Performance by category
| Category | Qwen 3.7 Plus average | MiniMax-M3 average |
|---|---|---|
| Legal | 8.17% | 16.99% |
| Finance | 48.11% | 53.52% |
| Education | 39.25% | 50.57% |
| Coding | 37.85% | 39.66% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Qwen 3.7 Plus cost | MiniMax-M3 cost | Qwen 3.7 Plus latency | MiniMax-M3 latency |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | $0.23 | $1.46 | 8m49s | 22m31s |
| Legal Research Bench | $0.30 | $0.34 | 13m52s | 13m34s |
| EMB | $0.60 | $2.09 | 36m49s | 31m30s |
| Finance Agent (v2) | $0.36 | $0.32 | 9m02s | 8m17s |
| MortgageTax | N/A | N/A | 60.55s | 26.76s |
| Tax Agent Bench | $0.09 | $0.16 | 5m10s | 5m15s |
| SAGE | N/A | N/A | 2m57s | 118.54s |
| Code Migration | $0.42 | $7.07 | 35m47s | 1h14m |
| SkillsBench | $0.14 | $0.50 | 12m04s | 15m28s |
| Vibe Code Bench v1.1 | $1.08 | $6.45 | 37m12s | 1h25m |
Results available only for Qwen 3.7 Plus
None.
Results available only for MiniMax-M3
- Vals Index
- LegalBench
- TaxEval v2
- MedCode
- MedScribe
- ProofBench v1.1
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- LiveCodeBench
- SWE-bench
- Terminal-Bench 4.0
- Vibe Code Bench 1-100
- CyberBench v1.1
- Public Benefits Bench v1.1