MiniMax-M3 vs GLM 5.3 Flash: Benchmark Comparison

MiniMax-M3 has the higher score on 5 of 19 shared benchmarks; GLM 5.3 Flash leads on 14.

The largest observed score gap is 17.00 pts on SWE-bench , where GLM 5.3 Flash leads.

Reported ±1 standard-error ranges overlap on 6 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark MiniMax-M3 GLM 5.3 Flash Gap Reported uncertainty
Harvey's Legal Agent Benchmark 4.17% ±1.65 6.67% ±1.65 2.50 pts Reported ±1 SE ranges overlap
Legal Research Bench 29.81% ±3.18 45.19% ±3.46 15.38 pts Reported ±1 SE ranges do not overlap
LegalBench 85.42% ±0.44 83.93% ±0.43 1.49 pts Reported ±1 SE ranges do not overlap
EMB 47.76% ±2.95 55.93% ±3.56 8.17 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 48.27% ±0.44 57.85% ±2.06 9.58 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 49.69% ±3.25 61.96% ±3.26 12.27 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.73% ±0.86 75.59% ±0.84 2.86 pts Reported ±1 SE ranges do not overlap
MedScribe 87.25% ±1.96 88.94% ±1.91 1.68 pts Reported ±1 SE ranges overlap
ProofBench v1.1 18.00% ±3.86 21.00% ±4.09 3.00 pts Reported ±1 SE ranges overlap
GPQA Diamond 92.68% ±1.44 86.36% ±2.13 6.31 pts Reported ±1 SE ranges do not overlap
MMLU Pro 84.22% ±0.36 86.06% ±0.34 1.84 pts Reported ±1 SE ranges do not overlap
MMMU Pro 81.16% ±0.94 86.01% ±0.83 4.86 pts Reported ±1 SE ranges do not overlap
Code Migration 19.93% ±3.94 20.52% ±4.21 0.59 pts Reported ±1 SE ranges overlap
LiveCodeBench 82.15% ±1.05 80.51% ±1.05 1.64 pts Reported ±1 SE ranges overlap
SkillsBench 51.50% ±4.50 40.16% ±4.43 11.34 pts Reported ±1 SE ranges do not overlap
SWE-bench 75.00% ±1.94 92.00% ±1.21 17.00 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 9.22% ±2.86 16.03% ±3.28 6.81 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 47.57% ±5.44 30.76% ±5.28 16.81 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 65.00% ±6.10 67.86% ±5.57 2.86 pts Reported ±1 SE ranges overlap

Performance by category

Category MiniMax-M3 average GLM 5.3 Flash average
Legal 39.80% 45.26%
Finance 54.61% 62.83%
Healthcare 87.25% 88.94%
Math 18.00% 21.00%
Academic 86.02% 86.14%
Coding 47.56% 46.66%
Cyber 65.00% 67.86%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark MiniMax-M3 cost GLM 5.3 Flash cost MiniMax-M3 latency GLM 5.3 Flash latency
Harvey's Legal Agent Benchmark $1.46 $0.57 22m31s 52m17s
Legal Research Bench $0.34 $0.12 13m34s 51m07s
LegalBench N/A N/A 8.21s 12.52s
EMB $2.09 $0.71 31m30s 1h15m
Finance Agent (v2) $0.32 $0.05 8m17s 14m51s
Tax Agent Bench $0.16 $0.14 5m15s 21m37s
TaxEval v2 N/A N/A 94.04s 53.06s
MedScribe N/A N/A 2m04s 86.98s
ProofBench v1.1 $0.42 $0.22 11m59s 1h11m
GPQA Diamond N/A N/A 4m39s 2m11s
MMLU Pro N/A N/A 41.39s 44.76s
MMMU Pro N/A N/A 68.35s 28.22s
Code Migration $7.07 $3.13 1h14m 3h26m
LiveCodeBench N/A N/A 6m07s 5m02s
SkillsBench $0.50 $0.04 15m28s 12m09s
SWE-bench $0.42 $0.02 12m06s 1h04m
Vibe Code Bench 1-100 $3.03 $0.46 26m10s 39m03s
Vibe Code Bench v1.1 $6.45 $2.35 1h25m 1h24m
CyberBench v1.1 $3.41 $0.14 40m56s 35m48s

Results available only for MiniMax-M3

  • Vals Index
  • MortgageTax
  • MedCode
  • SAGE
  • Terminal-Bench 4.0
  • Public Benefits Bench v1.1

Results available only for GLM 5.3 Flash

  • Terminal-Bench Science
  • IOI
  • ProgramBench
Model details MiniMax-M3 Model details GLM 5.3 Flash