Model comparison

Claude Sonnet 5 vs Grok 4.6: Benchmark Comparison

Claude Sonnet 5 has the higher score on 8 of 24 shared benchmarks; Grok 4.6 leads on 16.

The largest observed score gap is 26.00 pts on ProofBench v1.1 , where Claude Sonnet 5 leads.

Reported ±1 standard-error ranges overlap on 11 of 24 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 Grok 4.6 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 5.00% ±1.65 15.83% ±2.65 10.83 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.83% ±3.43 48.08% ±3.47 6.25 pts Reported ±1 SE ranges overlap
LegalBench 83.92% ±0.46 86.31% ±0.42 2.39 pts Reported ±1 SE ranges do not overlap
EMB 66.32% ±3.01 62.73% ±3.08 3.59 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 53.68% ±0.67 0.23 pts Reported ±1 SE ranges overlap
MortgageTax 70.03% ±0.90 64.19% ±0.95 5.84 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 70.79% ±3.00 8.51 pts Reported ±1 SE ranges do not overlap
TaxEval v2 75.63% ±0.84 71.10% ±0.90 4.54 pts Reported ±1 SE ranges do not overlap
MedCode 47.54% ±2.27 44.71% ±2.26 2.83 pts Reported ±1 SE ranges overlap
MedScribe 76.05% ±3.05 86.53% ±1.96 10.48 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 51.00% ±5.02 26.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 2.86% ±2.01 11.43% ±3.83 8.57 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 88.89% ±2.22 94.70% ±1.13 5.81 pts Reported ±1 SE ranges do not overlap
MMLU Pro 87.55% ±0.37 89.40% ±0.30 1.86 pts Reported ±1 SE ranges do not overlap
SAGE 48.92% ±3.40 28.90% ±3.08 20.02 pts Reported ±1 SE ranges do not overlap
Code Migration 44.39% ±4.25 44.57% ±4.48 0.18 pts Reported ±1 SE ranges overlap
IOI 45.00% ±2.75 47.61% ±2.09 2.61 pts Reported ±1 SE ranges overlap
LiveCodeBench 82.43% ±1.09 88.22% ±0.94 5.80 pts Reported ±1 SE ranges do not overlap
SkillsBench 46.48% ±4.49 55.77% ±4.84 9.29 pts Reported ±1 SE ranges overlap
SWE-bench 79.60% ±1.80 95.60% ±0.92 16.00 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 13.82% ±2.98 14.75% ±3.17 0.93 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 81.33% ±3.05 76.24% ±3.82 5.09 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 67.68% ±5.87 5.77 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.03% ±1.23 66.85% ±1.23 0.81 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Sonnet 5 average Grok 4.6 average
Legal 43.58% 50.07%
Finance 65.63% 64.50%
Healthcare 61.80% 65.62%
Math 77.00% 51.00%
Science 2.86% 11.43%
Academic 88.22% 92.05%
Education 48.92% 28.90%
Coding 56.15% 60.39%
Beta 61.91% 67.68%
Social Mobility 66.03% 66.85%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost Grok 4.6 cost Claude Sonnet 5 latency Grok 4.6 latency
Harvey's Legal Agent Benchmark $8.95 $4.01 38m52s 45m05s
Legal Research Bench $2.72 $1.53 25m46s 25m04s
LegalBench N/A N/A 4.90s 27.42s
EMB $10.29 $3.06 49m46s 52m59s
Finance Agent (v2) $0.75 $1.66 13m12s 16m41s
MortgageTax N/A N/A 28.67s 8.55s
Tax Agent Bench $1.79 $0.98 19m19s 10m58s
TaxEval v2 N/A N/A 3m22s 47.57s
MedCode N/A N/A 2m15s 2m07s
MedScribe N/A N/A 4m12s 70.94s
ProofBench v1.1 $1.37 $0.76 14m58s 14m22s
Terminal-Bench Science $18.70 $4.82 2h55m 1h18m
GPQA Diamond N/A N/A 63.23s 3m29s
MMLU Pro N/A N/A 25.19s 54.98s
SAGE N/A N/A 7m14s 3m22s
Code Migration $35.31 $14.58 1h57m 1h30m
IOI $12.89 $8.18 59m44s 3h57m
LiveCodeBench N/A N/A 77.02s 118.40s
SkillsBench $2.90 $0.92 15m12s 13m28s
SWE-bench $1.49 $0.78 16m02s 10m04s
Vibe Code Bench 1-100 $71.15 $26.27 4h16m 2h11m
Vibe Code Bench v1.1 $25.39 $4.88 1h07m 25m30s
CyberBench v1.1 $1.89 $4.07 16m57s 36m52s
Public Benefits Bench v1.1 $1.29 $0.92 21m09s 25m53s

Results available only for Claude Sonnet 5

  • Vals Index
  • MMMU Pro
  • ProgramBench
  • Terminal-Bench 4.0

Results available only for Grok 4.6

  • Vals RSI Index
  • BioMysteryBench
  • MysteryMechanism
  • Time Horizon Index: KSP
Model details Claude Sonnet 5 Model details Grok 4.6