Release Date: Sep 10, 2026

Developer DeepSeekย ๐Ÿ‡จ๐Ÿ‡ณ
Context Window 1M
Max Output Tokens 384k
Token Costs (in/out) $0.30/1.20
Weights Open
Input Modalities

Accuracy

57.86 % ยฑ 1.16

Cost / Test (Vals Index)

$ 0.303

Latency

25 min 34 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.16
15/56

0.0%

ยฑ4.29
9/58

0.0%

ยฑ3.23
25/56

0.0%

ยฑ0.39
22/59

0.0%

ยฑ3.42
14/59

0.0%

ยฑ2.04
46/91

0.0%

ยฑ1.92
19/93

0.0%

ยฑ5.01
13/30

0.0%

ยฑ3.44
26/81

0.0%

ยฑ1.25
13/34

0.0%

ยฑ0.54
7/9

0.0%

ยฑ3.15
12/19

0.0%

ยฑ2.89
7/94

0.0%

ยฑ0.46
51/143

0.0%

ยฑ3.88
1/34

0.0%

ยฑ1.63
14/64
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : DeepSeek
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 384,000
Reasoning Effort: high

Updates

Sep 10, 2026

We evaluated DeepSeekโ€™s DeepSeek V4.1 Flash across our benchmark suite.

  • It is the new #1 open-weight model on the Vals Index (57.86%, #15 of 56 overall), narrowly ahead of Kimi K3 (57.81%) at $0.30 per test versus $6.47, and finishing tasks in well under half the time. It is the cheapest model in the top 15.

  • It is the top open-weight model on Code Migration (45.62%, #9 of 58) at under $1 per test, while the next open-weight model, GLM 5.3 (44.22%), costs $24.91. It also takes #1 of 34 on SkillsBench (69.80% with skills, 61.66% without).

  • On Vibe Code Bench it is #2 among open-weight models (84.74%), 0.2 points behind Kimi K3 but roughly 40x cheaper ($0.41 vs $17.59 per task) and about 5x faster (15 minutes vs 1.5 hours). It is also the #2 open-weight model on Terminal-Bench 2.1 (74.53% across three full trials).

  • Compared to its predecessor, DeepSeek V4 Flash 0731, it gains 4.3 points on the Vals Index, with the largest jumps on Legal Research Bench (+11.1 points), Vibe Code Bench (+10.0) and Terminal-Bench 2.1 (+7.5). Despite list prices two to four times higher, it costs less per test on most agentic benchmarks because it finishes tasks in far fewer tokens, and it roughly halves latency on Vibe Code Bench, EMB, Legal Research and Harveyโ€™s Legal Agent Benchmark.

The model has a 1M-token context window and supports up to 384k output tokens. Evaluations were run at temperature 1 with default top-p and high reasoning effort.

Congrats to the DeepSeek team on another strong release!