Release Date: Dec 17, 2025

Developer Googleย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $0.50/3.00
Weights Private
Input Modalities

Accuracy

29.44 % ยฑ 0.98

Cost / Test (Vals Index)

$ 0.409

Latency

5 min 48 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ0.98
37/46

0.0%

ยฑ1.09
43/49

0.0%

ยฑ2.91
40/46

0.0%

ยฑ0.43
32/49

0.0%

ยฑ2.69
30/49

0.0%

ยฑ2.11
4/85

0.0%

ยฑ1.90
74/84

0.0%

ยฑ0.91
10/95

0.0%

ยฑ3.40
8/76

0.0%

ยฑ0.86
42/140

0.0%

ยฑ3.95
54/84

0.0%

ยฑ1.64
33/133

0.0%

ยฑ10.54
13/62

0.0%

ยฑ1.00
28/138

0.0%

ยฑ0.36
7/137

0.0%

ยฑ0.32
17/133

0.0%

ยฑ0.79
11/89

0.0%

ยฑ0.00
31/39

0.0%

ยฑ1.94
40/83

0.0%

ยฑ1.30
35/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Dec 18, 2025

Weโ€™ve run all our benchmarks for Gemini 3 Flash (12/25) and found it continues to excel!

Congrats again to Google on the release!

Dec 17, 2025

We evaluated Gemini 3 Flash (12/25) and found strong performance at a small fraction of a cost of other frontier models like GPT 5.2 and Claude Opus 4.5 (Thinking).

While the model doesnโ€™t achieve SOTA performance on any of our evaluations, it certainly comes close! It places second on SAGE , only 0.2% behind Claude Opus 4.5 (Thinking) for less than a tenth of the cost per token!

Incredibly, Gemini 3 Flash (12/25) outperforms Gemini 3 Pro (11/25) according to our comparison page

Congratulations to the Google team on another outstanding model!