Qwen 3.7 Plus vs GPT 5.4 Nano: Benchmark Comparison
Qwen 3.7 Plus has the higher score on 7 of 9 shared benchmarks; GPT 5.4 Nano leads on 1.
The largest observed score gap is 20.29 pts on Vibe Code Bench v1.1 , where Qwen 3.7 Plus leads.
Reported ±1 standard-error ranges overlap on 5 of 9 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Qwen 3.7 Plus | GPT 5.4 Nano | Gap | Reported uncertainty |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | 0.00% ±0.00 | 0.00% ±0.00 | 0.00 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 16.35% ±2.57 | 6.25% ±1.68 | 10.10 pts | Reported ±1 SE ranges do not overlap |
| EMB | 49.34% ±3.18 | 44.75% ±3.12 | 4.59 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 38.22% ±1.04 | 38.22% ±1.19 | 0.00 pts | Reported ±1 SE ranges overlap |
| MortgageTax | 66.18% ±0.93 | 59.10% ±0.96 | 7.07 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 38.71% ±2.81 | 26.58% ±2.52 | 12.14 pts | Reported ±1 SE ranges do not overlap |
| SAGE | 39.25% ±3.38 | 38.08% ±3.09 | 1.17 pts | Reported ±1 SE ranges overlap |
| Code Migration | 12.86% ±2.93 | 14.47% ±4.03 | 1.61 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 46.39% ±4.61 | 26.10% ±5.08 | 20.29 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Qwen 3.7 Plus average | GPT 5.4 Nano average |
|---|---|---|
| Legal | 8.17% | 3.13% |
| Finance | 48.11% | 42.16% |
| Education | 39.25% | 38.08% |
| Coding | 29.62% | 20.28% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Qwen 3.7 Plus cost | GPT 5.4 Nano cost | Qwen 3.7 Plus latency | GPT 5.4 Nano latency |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | $0.23 | $0.18 | 8m49s | 6m03s |
| Legal Research Bench | $0.30 | $0.14 | 13m52s | 8m47s |
| EMB | $0.60 | $0.57 | 36m49s | 29m27s |
| Finance Agent (v2) | $0.36 | $0.16 | 9m02s | 5m35s |
| MortgageTax | N/A | N/A | 60.55s | 14.53s |
| Tax Agent Bench | $0.09 | $0.07 | 5m10s | 3m40s |
| SAGE | N/A | N/A | 2m57s | 34.33s |
| Code Migration | $0.42 | $0.42 | 35m47s | 26m34s |
| Vibe Code Bench v1.1 | $1.08 | $1.28 | 37m12s | 54m36s |
Results available only for Qwen 3.7 Plus
- SkillsBench
Results available only for GPT 5.4 Nano
- LegalBench
- TaxEval v2
- MedCode
- MedScribe
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- LiveCodeBench
- SWE-bench