Release Date: Sep 25, 2025

Developer GoogleΒ πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $0.30/2.50
Weights Private
Input Modalities

Accuracy

64.88 %

Avg. Cost (In/Out)

$ 0.30 / $ 2.50

Latency

57.62 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.92
49/86

0.0%

Β±1.99
46/85

0.0%

Β±0.96
55/95

0.0%

Β±3.34
47/76

0.0%

Β±0.87
61/141

0.0%

Β±2.15
74/133

0.0%

Β±1.14
74/138

0.0%

Β±0.41
56/137

0.0%

Β±0.36
67/133

0.0%

Β±0.95
39/89
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,535

Updates

Sep 26, 2025

We evaluated Gemini 2.5 Flash Preview (9/25) (Thinking) (and also the Flash Lite model) and found the following:

  • Compared to the previous version, the update shows improvements on GPQA Diamond, with a ~17% increase over the previous version.

  • Flash improved on Terminal-Bench (+5%), GPQA Diamond (+17.2%), and our private Corp Fin Benchmark (+4.4%). It also ranks #3/38 on MMMU Pro and #6/20 on SWE-bench Verified (for the thinking model), while delivering performance at half the cost of similar models.

  • Flash Lite matches Flash on several public benchmarks, making it a very cost-effective option. However, Flash outperforms Lite by ~10% on our private benchmarks (CaseLaw v2, TaxEval, Mortgage Tax).

  • Flash delivers competitive performance at a fraction of the cost of other foundation models.

Overall, the latest update to Gemini 2.5 Flash is a highly efficient model that balances strong performance with low cost.