Model comparison

Qwen 3.8 Max vs Claude Sonnet 4.6: Benchmark Comparison

Qwen 3.8 Max has the higher score on 12 of 19 shared benchmarks; Claude Sonnet 4.6 leads on 7.

The largest observed score gap is 15.93 pts on Code Migration , where Claude Sonnet 4.6 leads.

Reported ±1 standard-error ranges overlap on 7 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Qwen 3.8 Max Claude Sonnet 4.6 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 10.42% ±2.29 5.00% ±1.65 5.42 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 47.60% ±3.47 38.46% ±3.38 9.13 pts Reported ±1 SE ranges do not overlap
LegalBench 83.61% ±0.42 82.12% ±0.48 1.49 pts Reported ±1 SE ranges do not overlap
EMB 60.07% ±3.35 60.15% ±3.28 0.09 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 50.59% ±0.58 51.03% ±0.33 0.44 pts Reported ±1 SE ranges overlap
MortgageTax 63.99% ±0.94 67.73% ±0.92 3.74 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 65.96% ±3.18 64.85% ±3.11 1.11 pts Reported ±1 SE ranges overlap
TaxEval v2 75.55% ±0.84 77.11% ±0.82 1.55 pts Reported ±1 SE ranges overlap
GPQA Diamond 93.69% ±1.22 85.61% ±2.41 8.08 pts Reported ±1 SE ranges do not overlap
MMLU Pro 88.60% ±0.31 87.34% ±0.43 1.26 pts Reported ±1 SE ranges do not overlap
MMMU Pro 88.03% ±0.78 83.58% ±0.89 4.45 pts Reported ±1 SE ranges do not overlap
SAGE 51.25% ±3.42 46.58% ±3.43 4.67 pts Reported ±1 SE ranges overlap
Code Migration 23.96% ±4.31 39.89% ±4.14 15.93 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 87.85% ±0.95 82.09% ±1.06 5.76 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
SkillsBench 42.01% ±4.43 49.05% ±4.64 7.04 pts Reported ±1 SE ranges overlap
SWE-bench 85.60% ±1.57 77.40% ±1.87 8.20 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 64.70% ±5.36 51.48% ±4.64 13.22 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 67.12% ±1.22 62.45% ±1.26 4.67 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Qwen 3.8 Max average Claude Sonnet 4.6 average
Legal 47.21% 41.86%
Finance 63.23% 64.17%
Academic 90.11% 85.51%
Education 51.25% 46.58%
Coding 50.69% 50.07%
Social Mobility 67.12% 62.45%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Qwen 3.8 Max cost Claude Sonnet 4.6 cost Qwen 3.8 Max latency Claude Sonnet 4.6 latency
Harvey's Legal Agent Benchmark $2.37 $3.04 49m12s 18m43s
Legal Research Bench $2.49 $2.30 1h01m 16m30s
LegalBench N/A N/A 36.37s 21.63s
EMB $2.66 $7.23 57m47s 54m16s
Finance Agent (v2) $1.24 $2.41 19m35s 11m36s
MortgageTax N/A N/A 47.50s 29.74s
Tax Agent Bench $1.55 $1.19 40m36s 11m47s
TaxEval v2 N/A N/A 4m26s 2m08s
GPQA Diamond N/A N/A 4m27s 6m18s
MMLU Pro N/A N/A 63.58s 55.41s
MMMU Pro N/A N/A 43.85s 2m12s
SAGE N/A N/A 2m13s 2m43s
Code Migration $10.54 $38.51 4h28m 1h40m
LiveCodeBench N/A N/A 5m38s 4m55s
ProgramBench N/A N/A 5h54m 1h12m
SkillsBench $0.52 $1.38 15m48s 14m20s
SWE-bench $1.13 $1.30 41m42s 8m32s
Vibe Code Bench v1.1 $8.24 $5.91 2h49m 26m12s
Public Benefits Bench v1.1 $0.99 $1.11 40m44s 14m43s

Results available only for Qwen 3.8 Max

  • Vals Index
  • Vals RSI Index
  • MedCode
  • MedScribe
  • ProofBench v1.1
  • MysteryMechanism
  • Terminal-Bench Science
  • IOI
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100
  • CyberBench v1.1

Results available only for Claude Sonnet 4.6

None.

Model details Qwen 3.8 Max Model details Claude Sonnet 4.6