DeepSeek V4.1 Flash vs GPT 5.5: Benchmark Comparison
DeepSeek V4.1 Flash has the higher score on 8 of 15 shared benchmarks; GPT 5.5 leads on 6.
The largest observed score gap is 14.89 pts on Vibe Code Bench v1.1 , where DeepSeek V4.1 Flash leads.
Reported ±1 standard-error ranges overlap on 8 of 15 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | DeepSeek V4.1 Flash | GPT 5.5 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | 6.67% ±1.83 | 3.75% ±1.17 | 2.92 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 41.35% ±3.42 | 40.38% ±3.41 | 0.96 pts | Reported ±1 SE ranges overlap |
| LegalBench | 83.28% ±0.46 | 86.52% ±0.41 | 3.24 pts | Reported ±1 SE ranges do not overlap |
| EMB | 57.21% ±3.23 | 64.54% ±2.87 | 7.34 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 53.48% ±0.39 | 51.76% ±0.55 | 1.72 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 62.46% ±3.15 | 60.46% ±3.22 | 2.00 pts | Reported ±1 SE ranges overlap |
| MedCode | 41.17% ±2.04 | 49.10% ±2.19 | 7.93 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 85.50% ±1.92 | 86.87% ±1.93 | 1.37 pts | Reported ±1 SE ranges overlap |
| SAGE | 47.88% ±3.44 | 51.53% ±3.95 | 3.65 pts | Reported ±1 SE ranges overlap |
| Code Migration | 45.62% ±4.29 | 45.16% ±4.16 | 0.46 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 0.50% ±0.50 | 0.50% ±0.50 | 0.00 pts | Reported ±1 SE ranges overlap |
| SkillsBench | 69.80% ±3.88 | 62.21% ±4.52 | 7.59 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 84.74% ±2.89 | 69.85% ±4.54 | 14.89 pts | Reported ±1 SE ranges do not overlap |
| SRE Bench | 0.76% ±0.54 | 3.82% ±1.19 | 3.05 pts | Reported ±1 SE ranges do not overlap |
| Public Benefits Bench v1.1 | 64.28% ±1.25 | 60.89% ±1.27 | 3.38 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | DeepSeek V4.1 Flash average | GPT 5.5 average |
|---|---|---|
| Legal | 43.76% | 43.55% |
| Finance | 57.72% | 58.92% |
| Healthcare | 63.34% | 67.98% |
| Education | 47.88% | 51.53% |
| Coding | 50.17% | 44.43% |
| Cyber | 0.76% | 3.82% |
| Social Mobility | 64.28% | 60.89% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | DeepSeek V4.1 Flash cost | GPT 5.5 cost | DeepSeek V4.1 Flash latency | GPT 5.5 latency |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | $0.17 | $4.60 | 9m39s | 12m14s |
| Legal Research Bench | $0.25 | $7.40 | 11m42s | 37m45s |
| LegalBench | N/A | N/A | 6.53s | 18.14s |
| EMB | $0.23 | $3.27 | 13m00s | 15m11s |
| Finance Agent (v2) | $0.21 | $4.15 | 6m04s | 11m02s |
| Tax Agent Bench | $0.13 | $4.07 | 5m49s | 15m52s |
| MedCode | N/A | N/A | 35.13s | 2m40s |
| MedScribe | N/A | N/A | 53.83s | 2m13s |
| SAGE | N/A | N/A | 35.78s | 76.05s |
| Code Migration | $0.94 | $6.44 | 1h45m | 31m35s |
| ProgramBench | $0.90 | $6.95 | 1h02m | 22m03s |
| SkillsBench | $0.09 | $2.54 | 3m43s | 6m55s |
| Vibe Code Bench v1.1 | $0.41 | $16.66 | 15m28s | 31m52s |
| SRE Bench | $0.55 | $17.77 | 1h01m | 46m28s |
| Public Benefits Bench v1.1 | $0.07 | $3.97 | 9m59s | 39m37s |
Results available only for DeepSeek V4.1 Flash
- Vals Index
- ProofBench v1.1
- BioMysteryBench
- MysteryMechanism
- Terminal-Bench Science
- IOI
- Terminal-Bench 4.0
- Vibe Code Bench 1-100
- CyberBench v1.1
Results available only for GPT 5.5
- Vals RSI Index
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- LiveCodeBench
- SWE-bench
- Time Horizon Index: KSP