Model comparison

Claude Sonnet 5 vs Grok 4.7: Benchmark Comparison

Claude Sonnet 5 has the higher score on 4 of 19 shared benchmarks; Grok 4.7 leads on 15.

The largest observed score gap is 51.00 pts on ProofBench v1.1 , where Claude Sonnet 5 leads.

Reported ±1 standard-error ranges overlap on 10 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 Grok 4.7 Gap Reported uncertainty
Vals Index 51.77% ±1.09 54.95% ±1.07 3.17 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 12.50% ±2.42 7.50 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.83% ±3.43 47.12% ±3.47 5.29 pts Reported ±1 SE ranges overlap
LegalBench 83.92% ±0.46 84.39% ±0.46 0.47 pts Reported ±1 SE ranges overlap
EMB 66.32% ±3.01 66.99% ±3.04 0.67 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 52.25% ±0.35 1.66 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 65.60% ±3.24 3.32 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 49.55% ±2.17 2.01 pts Reported ±1 SE ranges overlap
MedScribe 76.05% ±3.05 89.38% ±1.89 13.32 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 26.00% ±4.41 51.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 2.86% ±2.01 10.00% ±3.61 7.14 pts Reported ±1 SE ranges do not overlap
SAGE 48.92% ±3.40 40.79% ±3.35 8.13 pts Reported ±1 SE ranges do not overlap
Code Migration 44.39% ±4.25 44.82% ±4.21 0.42 pts Reported ±1 SE ranges overlap
IOI 45.00% ±2.75 57.72% ±1.99 12.72 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 9.60% ±1.01 28.79% ±2.31 19.19 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 81.33% ±3.05 86.17% ±2.18 4.85 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 69.46% ±5.67 7.56 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.03% ±1.23 65.63% ±1.24 0.41 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Sonnet 5 average Grok 4.7 average
Legal 43.58% 48.00%
Finance 60.83% 61.61%
Healthcare 61.80% 69.47%
Math 77.00% 26.00%
Science 2.86% 10.00%
Education 48.92% 40.79%
Coding 36.06% 43.60%
Beta 61.91% 69.46%
Social Mobility 66.03% 65.63%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost Grok 4.7 cost Claude Sonnet 5 latency Grok 4.7 latency
Vals Index $13.72 $12.12 54m16s 35m31s
Harvey's Legal Agent Benchmark $8.95 $11.13 38m52s 42m05s
Legal Research Bench $2.72 $4.92 25m46s 21m14s
LegalBench N/A N/A 4.90s 26.02s
EMB $10.29 $6.48 49m46s 34m57s
Finance Agent (v2) $0.75 $2.72 13m12s 15m33s
Tax Agent Bench $1.79 $1.80 19m19s 16m36s
MedCode N/A N/A 2m15s 3m33s
MedScribe N/A N/A 4m12s 2m33s
ProofBench v1.1 $1.37 $0.79 14m58s 12m47s
Terminal-Bench Science $18.70 $14.45 2h55m 1h29m
SAGE N/A N/A 7m14s 4m51s
Code Migration $35.31 $36.55 1h57m 1h09m
IOI $12.89 $12.71 59m44s 46m48s
ProgramBench $24.36 $46.49 1h31m 3h38m
Terminal-Bench 4.0 $26.33 $18.09 1h45m 50m21s
Vibe Code Bench v1.1 $25.39 $15.83 1h07m 35m47s
CyberBench v1.1 $1.89 $8.54 16m57s 15m30s
Public Benefits Bench v1.1 $1.29 $2.25 21m09s 16m57s

Results available only for Claude Sonnet 5

  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for Grok 4.7

  • Vals RSI Index
  • BioMysteryBench
  • MysteryMechanism
Model details Claude Sonnet 5 Model details Grok 4.7