Release Date: Jul 21, 2026

Developer Googleย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $1.50/7.50
Weights Private
Input Modalities

Accuracy

55.35 % ยฑ 1.09

Cost / Test (Vals Index)

$ 3.030

Latency

22 min 8 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.09
14/46

0.0%

ยฑ4.07
19/49

0.0%

ยฑ4.51
18/22

0.0%

ยฑ2.65
10/46

0.0%

ยฑ0.18
7/49

0.0%

ยฑ3.01
27/49

0.0%

ยฑ2.16
9/85

0.0%

ยฑ1.86
42/84

0.0%

ยฑ0.92
23/95

0.0%

ยฑ3.38
22/76

0.0%

ยฑ1.29
22/30

0.0%

ยฑ0.85
24/140

0.0%

ยฑ4.29
22/84

0.0%

ยฑ1.33
6/133

0.0%

ยฑ0.94
6/138

0.0%

ยฑ0.41
8/137

0.0%

ยฑ0.30
11/133

0.0%

ยฑ0.77
5/89

0.0%

ยฑ0.00
21/39

0.0%

ยฑ1.80
21/83

0.0%

ยฑ1.63
11/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Jul 21, 2026

We evaluated Googleโ€™s Gemini 3.6 Flash on the Vals Index and proprietary benchmarks.

  • Gemini 3.6 Flash scores 55.35% on the Vals Index, placing #13 of 43 models overall and finishing within 2.28 points of Gemini 3.5 Flash.

  • Gemini 3.6 Flash ranks #1 on the CyberBench Patch track at 84.75%, which measures fixing real-world open-source vulnerabilities.

  • Its coding results include 77.45% on the SWE-bench Verified Vals Index subset and 57.88% on the Vibe Code Bench Vals Index subset, improvements of 1.96 and 3.17 points over Gemini 3.5 Flash, respectively.

  • Gemini 3.6 Flash scores 73.78% across three full trials of Terminal-Bench 2.1, placing #8 of 43 models. It also places #7 of 74 models on MedCode, scoring 53.15%.

We evaluated Gemini 3.6 Flash with temperature 1, high reasoning effort, and up to 65k output tokens. The model supports a 1M-token context window, multimodal inputs, and tool calling.

Congrats to the Google team on the release!