DeepSeek V4.1 Flash vs GPT-6 Luna: Benchmark Comparison

DeepSeek V4.1 Flash has the higher score on 12 of 20 shared benchmarks; GPT-6 Luna leads on 7.

The largest observed score gap is 15.28 pts on IOI , where GPT-6 Luna leads.

Reported ±1 standard-error ranges overlap on 10 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4.1 Flash GPT-6 Luna Gap Reported uncertainty
Vals Index 51.32% ±1.13 51.22% ±1.07 0.10 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 6.67% ±1.83 2.92% ±1.36 3.75 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.35% ±3.42 30.29% ±3.19 11.06 pts Reported ±1 SE ranges do not overlap
EMB 57.21% ±3.23 68.52% ±2.81 11.31 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 53.48% ±0.39 49.87% ±0.23 3.61 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.46% ±3.15 58.86% ±3.21 3.59 pts Reported ±1 SE ranges overlap
MedCode 41.17% ±2.04 44.69% ±2.30 3.51 pts Reported ±1 SE ranges overlap
MedScribe 85.50% ±1.92 83.71% ±1.95 1.79 pts Reported ±1 SE ranges overlap
ProofBench v1.1 54.00% ±5.01 64.00% ±4.82 10.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 67.78% ±1.11 61.48% ±2.59 6.30 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 21.17% ±2.75 19.37% ±2.66 1.80 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 0.00% ±0.00 4.29% ±2.44 4.29 pts Reported ±1 SE ranges do not overlap
SAGE 47.88% ±3.44 48.09% ±3.38 0.21 pts Reported ±1 SE ranges overlap
Code Migration 45.62% ±4.29 42.55% ±4.42 3.07 pts Reported ±1 SE ranges overlap
IOI 40.28% ±2.56 55.56% ±8.87 15.28 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 19.70% ±1.75 13.64% ±1.51 6.06 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 84.74% ±2.89 81.65% ±3.38 3.09 pts Reported ±1 SE ranges overlap
CyberBench v1.1 73.69% ±5.48 76.25% ±5.29 2.56 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 64.28% ±1.25 57.65% ±1.28 6.63 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4.1 Flash average GPT-6 Luna average
Legal 24.01% 16.60%
Finance 57.72% 59.08%
Healthcare 63.34% 64.20%
Math 54.00% 64.00%
Science 29.65% 28.38%
Education 47.88% 48.09%
Coding 38.17% 38.78%
Cyber 73.69% 76.25%
Social Mobility 64.28% 57.65%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4.1 Flash cost GPT-6 Luna cost DeepSeek V4.1 Flash latency GPT-6 Luna latency
Vals Index $0.33 $0.43 27m53s 30m34s
Harvey's Legal Agent Benchmark $0.17 $0.30 9m39s 17m32s
Legal Research Bench $0.25 $0.44 11m42s 42m28s
EMB $0.23 $0.14 13m00s 17m21s
Finance Agent (v2) $0.21 $0.12 6m04s 14m58s
Tax Agent Bench $0.13 $0.19 5m49s 33m16s
MedCode N/A N/A 35.13s 108.77s
MedScribe N/A N/A 53.83s 2m50s
ProofBench v1.1 $0.13 $0.04 9m07s 7m31s
BioMysteryBench $0.23 $0.05 14m50s 8m44s
MysteryMechanism $0.16 $0.03 8m04s 6m01s
Terminal-Bench Science $0.47 $0.24 2h57m 1h16m
SAGE N/A N/A 35.78s 91.07s
Code Migration $0.94 $0.60 1h45m 50m23s
IOI $0.24 $0.20 18m30s 35m47s
ProgramBench $0.90 $0.18 1h02m 28m55s
Terminal-Bench 4.0 $0.50 $0.35 48m22s 32m46s
Vibe Code Bench v1.1 $0.41 $1.35 15m28s 37m00s
CyberBench v1.1 $0.07 $0.13 13m50s 15m50s
Public Benefits Bench v1.1 $0.07 $0.22 9m59s 52m59s

Results available only for DeepSeek V4.1 Flash

  • LegalBench
  • SkillsBench
  • Vibe Code Bench 1-100
  • SRE Bench

Results available only for GPT-6 Luna

None.

Model details DeepSeek V4.1 Flash Model details GPT-6 Luna