Release Date: Mar 3, 2026

Developer Googleย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $0.25/1.50
Weights Private
Input Modalities

Accuracy

15.46 % ยฑ 0.42

Cost / Test (Vals Index)

$ 0.136

Latency

2 min 43 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ0.42
47/47

0.0%

ยฑ1.07
47/49

0.0%

ยฑ1.09
47/47

0.0%

ยฑ0.64
46/50

0.0%

ยฑ1.25
47/50

0.0%

ยฑ2.07
22/86

0.0%

ยฑ1.82
81/85

0.0%

ยฑ0.91
19/95

0.0%

ยฑ3.48
17/76

0.0%

ยฑ0.88
70/141

0.0%

ยฑ0.00
83/85

0.0%

ยฑ1.97
61/133

0.0%

ยฑ1.08
66/138

0.0%

ยฑ0.41
40/137

0.0%

ยฑ0.34
42/133

0.0%

ยฑ0.91
33/89

0.0%

ยฑ0.00
38/39

0.0%

ยฑ2.16
72/84

0.0%

ยฑ0.99
54/56
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Mar 3, 2026

We evaluated Gemini 3.1 Flash Lite Preview across our full benchmark suite. The model is Googleโ€™s fast and cost-efficient offering.

  • The model does well on two of our multimodal benchmarks, SAGE (ranked 5th, with 49.5% accuracy) and Mortgage Tax (ranked 7th, 67.8% accuracy).
  • It ranks 14th on MMLU Pro with 86.2% accuracy.
  • On coding tasks the model still has room for improvementโ€”it ranks 29th on both Live Code Bench and Terminal-Bench 2.0, and 28th on SWE-bench Verified. It scores 0% on Vibe Code Bench.
  • Overall, it places 15th/20 on the Vals Multimodal Index and 22nd/31 on the Vals Index.
  • The cost savings compared to other models in the Gemini 3 series or Gemini 2.5 are dramatic: roughly 5โ€“20x cheaper per test across benchmarks, while maintaining respectable accuracy. For example, on Finance Agent it costs $0.072 per test vs $0.370 for Gemini 3 Flash, while performing comparably.

Our results show that the model does not perform as well as other models in the Gemini 3 series. However, it is fast and quite cost-efficient relative to those models, making it a good choice for applications that demand scale, speed, or cost-efficiency.

Evaluations were run with a temperature of 1.0 and a โ€œhighโ€ thinking level, via the official Google API.