Qwen 3.8 Max vs DeepSeek V4 Pro 0813: Benchmark Comparison
Qwen 3.8 Max has the higher score on 16 of 24 shared benchmarks; DeepSeek V4 Pro 0813 leads on 7.
The largest observed score gap is 44.28 pts on CyberBench v1.1 , where DeepSeek V4 Pro 0813 leads.
Reported ±1 standard-error ranges overlap on 10 of 24 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Qwen 3.8 Max | DeepSeek V4 Pro 0813 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 48.27% ±1.34 | 47.63% ±1.12 | 0.64 pts | Reported ±1 SE ranges overlap |
| Harvey's Legal Agent Benchmark | 10.42% ±2.29 | 7.50% ±2.15 | 2.92 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 47.60% ±3.47 | 40.87% ±3.42 | 6.73 pts | Reported ±1 SE ranges overlap |
| LegalBench | 83.61% ±0.42 | 82.36% ±0.44 | 1.25 pts | Reported ±1 SE ranges do not overlap |
| EMB | 60.07% ±3.35 | 52.80% ±3.06 | 7.26 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 50.59% ±0.58 | 50.39% ±0.25 | 0.20 pts | Reported ±1 SE ranges overlap |
| Tax Agent Bench | 65.96% ±3.18 | 58.66% ±3.14 | 7.30 pts | Reported ±1 SE ranges do not overlap |
| TaxEval v2 | 75.55% ±0.84 | 73.06% ±0.87 | 2.49 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 40.67% ±2.03 | 42.47% ±2.16 | 1.80 pts | Reported ±1 SE ranges overlap |
| MedScribe | 84.95% ±2.00 | 80.17% ±2.00 | 4.77 pts | Reported ±1 SE ranges do not overlap |
| ProofBench v1.1 | 58.00% ±4.96 | 50.00% ±5.03 | 8.00 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench Science | 12.86% ±4.03 | 4.29% ±2.44 | 8.57 pts | Reported ±1 SE ranges do not overlap |
| GPQA Diamond | 93.69% ±1.22 | 92.42% ±2.02 | 1.26 pts | Reported ±1 SE ranges overlap |
| MMLU Pro | 88.60% ±0.31 | 86.97% ±0.34 | 1.63 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 23.96% ±4.31 | 41.54% ±4.30 | 17.58 pts | Reported ±1 SE ranges do not overlap |
| IOI | 68.89% ±5.20 | 51.61% ±2.40 | 17.28 pts | Reported ±1 SE ranges do not overlap |
| LiveCodeBench | 87.85% ±0.95 | 87.53% ±0.96 | 0.33 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 0.00% ±0.00 | 0.00% ±0.00 | 0.00 pts | Reported ±1 SE ranges overlap |
| SkillsBench | 42.01% ±4.43 | 53.83% ±4.49 | 11.82 pts | Reported ±1 SE ranges do not overlap |
| SWE-bench | 85.60% ±1.57 | 96.40% ±0.83 | 10.80 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 34.34% ±3.94 | 14.14% ±2.02 | 20.20 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench 1-100 | 12.83% ±3.18 | 17.53% ±3.41 | 4.70 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 64.70% ±5.36 | 82.30% ±3.15 | 17.60 pts | Reported ±1 SE ranges do not overlap |
| CyberBench v1.1 | 28.57% ±3.31 | 72.86% ±5.50 | 44.28 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Qwen 3.8 Max average | DeepSeek V4 Pro 0813 average |
|---|---|---|
| Legal | 47.21% | 43.58% |
| Finance | 63.04% | 58.73% |
| Healthcare | 62.81% | 61.32% |
| Math | 58.00% | 50.00% |
| Science | 12.86% | 4.29% |
| Academic | 91.14% | 89.70% |
| Coding | 46.69% | 49.43% |
| Cyber | 28.57% | 72.86% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Qwen 3.8 Max cost | DeepSeek V4 Pro 0813 cost | Qwen 3.8 Max latency | DeepSeek V4 Pro 0813 latency |
|---|---|---|---|---|
| Vals Index | $4.96 | $3.35 | 1h41m | 1h04m |
| Harvey's Legal Agent Benchmark | $2.37 | $0.17 | 49m12s | 21m53s |
| Legal Research Bench | $2.49 | $1.09 | 1h01m | 42m07s |
| LegalBench | N/A | N/A | 36.37s | 20.05s |
| EMB | $2.66 | $1.08 | 57m47s | 37m27s |
| Finance Agent (v2) | $1.24 | $0.88 | 19m35s | 16m15s |
| Tax Agent Bench | $1.55 | $0.73 | 40m36s | 26m37s |
| TaxEval v2 | N/A | N/A | 4m26s | 3m46s |
| MedCode | N/A | N/A | 6m18s | 3m01s |
| MedScribe | N/A | N/A | 4m11s | 2m36s |
| ProofBench v1.1 | $1.50 | $0.06 | 31m08s | 12m12s |
| Terminal-Bench Science | $16.22 | $4.63 | 3h20m | 2h16m |
| GPQA Diamond | N/A | N/A | 4m27s | 4m39s |
| MMLU Pro | N/A | N/A | 63.58s | 59.62s |
| Code Migration | $10.54 | $18.58 | 4h28m | 3h32m |
| IOI | $9.47 | $2.15 | 1h47m | 1h07m |
| LiveCodeBench | N/A | N/A | 5m38s | 4m33s |
| ProgramBench | $12.36 | $0.43 | 5h54m | 1h23m |
| SkillsBench | $0.52 | $0.37 | 15m48s | 11m15s |
| SWE-bench | $1.13 | $0.10 | 41m42s | 4m00s |
| Terminal-Bench 4.0 | $10.63 | $3.31 | 2h28m | 1h32m |
| Vibe Code Bench 1-100 | $9.68 | $5.89 | 1h36m | 37m42s |
| Vibe Code Bench v1.1 | $8.24 | $0.36 | 2h49m | 1h06m |
| CyberBench v1.1 | $0.78 | $0.59 | 9m40s | 22m33s |
Results available only for Qwen 3.8 Max
- Vals RSI Index
- MortgageTax
- MysteryMechanism
- MMMU Pro
- SAGE
- Public Benefits Bench v1.1
Results available only for DeepSeek V4 Pro 0813
None.