Release Date: Apr 2, 2026

Developer GoogleΒ πŸ‡ΊπŸ‡Έ
Context Window 262k
Max Output Tokens 33k
Token Costs (in/out) $0.00/0.00
Weights Open
Input Modalities

Accuracy

51.91 %

Avg. Cost (In/Out)

$ 0.00 / $ 0.00

Latency

5 min 30 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.97
58/96

0.0%

Β±3.41
2/77
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 32,768
Reasoning Effort: high

Updates

Apr 9, 2026

We evaluated Gemma 4 31B IT β€” a 31B dense open model in Google DeepMind’s Gemma 4 family, with a 262k context window and 33k max output.

Key takeaways:

  • Its best result is #1 on SAGE at 55.03%, which makes Gemma 4 31B IT look strongest on education and exam-style reasoning.
  • It is also competitive on structured document and finance work: 61.37% on MortgageTax, 52.63% on Case Law (v2), and 50.79% on Finance Agent (v1.1).
  • On the subset tasks used in Vals Index, it posted 59.67% on the Corp Fin (v2) shared-max-context split and 53.92% on the SWE-bench Verified subset.
  • Its latency is strongest on structured tasks like Case Law (28.7s) and MortgageTax (41.9s), then rises sharply on heavier agentic workloads like Vals Index (323.2s), Terminal-Bench 2.0 (855.3s), and Finance Agent (5360.3s).
  • Overall, Gemma lands at 38.94% on Vals Index and 45.12% on Vals Multimodal Index. The main story here is efficiency for size, not broad frontier dominance.