Release Date: Feb 19, 2026

Developer GoogleΒ πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $2.00/12.00
Weights Private
Input Modalities

Accuracy

41.90 % Β± 1.17

Cost / Test (Vals Index)

$ 1.945

Latency

9 min 20 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.17
28/47

0.0%

Β±3.93
29/49

0.0%

Β±4.17
22/22

0.0%

Β±2.97
28/47

0.0%

Β±1.21
32/50

0.0%

Β±2.81
30/50

0.0%

Β±2.00
2/86

0.0%

Β±1.92
58/85

0.0%

Β±0.91
5/95

0.0%

Β±4.39
16/23

0.0%

Β±3.29
23/76

0.0%

Β±0.86
55/141

0.0%

Β±4.34
46/85

0.0%

Β±1.05
1/133

0.0%

Β±0.93
4/138

0.0%

Β±0.33
2/137

0.0%

Β±0.28
3/133

0.0%

Β±0.78
8/89

0.0%

Β±0.00
25/39

0.0%

Β±1.83
25/84

0.0%

Β±1.12
15/56
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Feb 20, 2026

We evaluated Gemini 3.1 Pro Preview (02/26) across our full benchmark suite. Here are the key takeaways:

One notable metric to call out here is the model achieves this performance at a lower cost than models like Claude Opus 4.6, Claude Sonnet 4.6, GPT 5.2 and O3.

Evaluations were run with a temperature of 1.0 and a β€œhigh” thinking level, via the official Google API.

Congratulations to the Google team on another outstanding model!

Feb 19, 2026

We evaluated Gemini 3.1 Pro Preview (02/26) on our Vals Index. Here are the key takeaways:

We are evaluating Gemini 3.1 Pro Preview (02/26) on our full suite of benchmarks and will be sharing updates soon!