Qwen 3.7 Plus vs GPT 5.4 Mini: Benchmark Comparison

Qwen 3.7 Plus has the higher score on 4 of 9 shared benchmarks; GPT 5.4 Mini leads on 4.

The largest observed score gap is 11.56 pts on SAGE , where GPT 5.4 Mini leads.

Reported ±1 standard-error ranges overlap on 6 of 9 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Qwen 3.7 Plus GPT 5.4 Mini Gap Reported uncertainty
Harvey's Legal Agent Benchmark 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
Legal Research Bench 16.35% ±2.57 12.50% ±2.30 3.85 pts Reported ±1 SE ranges overlap
EMB 49.34% ±3.18 45.43% ±3.65 3.92 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 38.22% ±1.04 45.36% ±0.45 7.14 pts Reported ±1 SE ranges do not overlap
MortgageTax 66.18% ±0.93 63.51% ±0.91 2.66 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 38.71% ±2.81 36.59% ±3.10 2.13 pts Reported ±1 SE ranges overlap
SAGE 39.25% ±3.38 50.81% ±3.40 11.56 pts Reported ±1 SE ranges do not overlap
Code Migration 12.86% ±2.93 12.94% ±3.64 0.08 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 46.39% ±4.61 47.97% ±5.61 1.58 pts Reported ±1 SE ranges overlap

Performance by category

Category Qwen 3.7 Plus average GPT 5.4 Mini average
Legal 8.17% 6.25%
Finance 48.11% 47.72%
Education 39.25% 50.81%
Coding 29.62% 30.45%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Qwen 3.7 Plus cost GPT 5.4 Mini cost Qwen 3.7 Plus latency GPT 5.4 Mini latency
Harvey's Legal Agent Benchmark $0.23 $1.00 8m49s 11m21s
Legal Research Bench $0.30 $1.86 13m52s 37m36s
EMB $0.60 $2.15 36m49s 43m21s
Finance Agent (v2) $0.36 $1.20 9m02s 35m54s
MortgageTax N/A N/A 60.55s 2m06s
Tax Agent Bench $0.09 $0.81 5m10s 16m46s
SAGE N/A N/A 2m57s 2m21s
Code Migration $0.42 $1.95 35m47s 28m53s
Vibe Code Bench v1.1 $1.08 $1.19 37m12s 34m15s

Results available only for Qwen 3.7 Plus

  • SkillsBench

Results available only for GPT 5.4 Mini

  • Vals Index
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • ProgramBench
  • SWE-bench
  • Terminal-Bench 4.0
Model details Qwen 3.7 Plus Model details GPT 5.4 Mini