Release Date: Jul 21, 2026

Developer Googleย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $0.30/2.50
Weights Private
Input Modalities

Accuracy

36.71 % ยฑ 1.06

Cost / Test (Vals Index)

$ 0.543

Latency

7 min 41 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.06
32/46

0.0%

ยฑ1.71
44/49

0.0%

ยฑ5.74
10/22

0.0%

ยฑ3.10
33/46

0.0%

ยฑ0.51
25/49

0.0%

ยฑ2.41
38/49

0.0%

ยฑ1.95
31/85

0.0%

ยฑ2.03
72/84

0.0%

ยฑ0.91
11/95

0.0%

ยฑ3.38
26/76

0.0%

ยฑ1.30
25/30

0.0%

ยฑ0.88
58/140

0.0%

ยฑ4.63
43/84

0.0%

ยฑ1.85
51/133

0.0%

ยฑ1.17
17/62

0.0%

ยฑ1.10
70/138

0.0%

ยฑ0.44
33/137

0.0%

ยฑ0.34
50/133

0.0%

ยฑ0.89
26/89

0.0%

ยฑ0.00
32/39

0.0%

ยฑ1.94
41/83

0.0%

ยฑ0.99
41/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Jul 21, 2026

We evaluated Googleโ€™s Gemini 3.5 Flash Lite on the Vals Index and across our benchmark suite.

We evaluated Gemini 3.5 Flash Lite with temperature 1, high reasoning effort, and up to 65k output tokens. The model supports a 1M-token context window, multimodal inputs, and tool calling.

Results are now available across 25 benchmarks. On the full 500-task SWE-bench Verified evaluation, Gemini 3.5 Flash Lite scores 75.00%; the Vals Index uses the 68.63% subset result above.