Release Date: May 19, 2026

Developer GoogleΒ πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $1.50/9.00
Weights Private
Input Modalities

Accuracy

53.08 % Β± 1.12

Cost / Test (Vals Index)

$ 2.915

Latency

11 min 1 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.12
18/47

0.0%

Β±4.10
22/49

0.0%

Β±5.79
11/22

0.0%

Β±2.75
13/47

0.0%

Β±0.23
4/50

0.0%

Β±3.21
21/50

0.0%

Β±2.11
5/86

0.0%

Β±1.92
56/85

0.0%

Β±0.91
18/95

0.0%

Β±4.63
14/23

0.0%

Β±3.44
15/76

0.0%

Β±0.85
34/141

0.0%

Β±4.73
34/85

0.0%

Β±1.46
12/133

0.0%

Β±0.95
10/138

0.0%

Β±0.85
43/137

0.0%

Β±0.31
8/133

0.0%

Β±0.77
6/89

0.0%

Β±0.00
23/39

0.0%

Β±4.47
15/29

0.0%

Β±1.83
24/84

0.0%

Β±1.12
10/56
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

May 19, 2026

The model has a 1M-token context window. Evaluations were run using a reasoning effort of β€œhigh”, a temperature of 1.0, and max output tokens set to 65k.

Congrats to the Google team on the strong release!