Release Date: Feb 19, 2026

Developer GoogleΒ πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $2.00/12.00
Weights Private
Input Modalities

Accuracy

41.90 % Β± 1.17

Cost / Test (Vals Index)

$ 1.945

Latency

9 min 20 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.17
36/57

0.0%

Β±3.93
36/59

0.0%

Β±4.17
25/25

0.0%

Β±2.97
36/57

0.0%

Β±1.21
40/60

0.0%

Β±2.81
38/60

0.0%

Β±2.00
2/92

0.0%

Β±1.92
64/94

0.0%

Β±0.91
6/98

0.0%

Β±4.39
21/31

0.0%

Β±3.29
24/81

0.0%

Β±0.86
58/145

0.0%

Β±4.34
53/95

0.0%

Β±1.05
1/138

0.0%

Β±0.93
6/143

0.0%

Β±0.33
3/144

0.0%

Β±0.28
4/138

0.0%

Β±0.78
10/93

0.0%

Β±0.00
30/45

0.0%

Β±1.83
29/88

0.0%

Β±1.12
21/65
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Feb 20, 2026

We evaluated Gemini 3.1 Pro Preview (02/26) across our full benchmark suite. Here are the key takeaways:

One notable metric to call out here is the model achieves this performance at a lower cost than models like Claude Opus 4.6, Claude Sonnet 4.6, GPT 5.2 and O3.

Evaluations were run with a temperature of 1.0 and a β€œhigh” thinking level, via the official Google API.

Congratulations to the Google team on another outstanding model!

Feb 19, 2026

We evaluated Gemini 3.1 Pro Preview (02/26) on our Vals Index. Here are the key takeaways:

We are evaluating Gemini 3.1 Pro Preview (02/26) on our full suite of benchmarks and will be sharing updates soon!