Apr 9, 2026
Gemma 4 31B IT is a size-efficient open release
We evaluated Gemma 4 31B IT β a 31B dense open model in Google DeepMindβs Gemma 4 family, with a 262k context window and 33k max output.
Key takeaways:
- Its best result is #1 on SAGE at 55.03%, which makes Gemma 4 31B IT look strongest on education and exam-style reasoning.
- It is also competitive on structured document and finance work: 61.37% on MortgageTax, 52.63% on Case Law (v2), and 50.79% on Finance Agent (v1.1).
- On the subset tasks used in Vals Index, it posted 59.67% on the Corp Fin (v2) shared-max-context split and 53.92% on the SWE-bench Verified subset.
- Its latency is strongest on structured tasks like Case Law (28.7s) and MortgageTax (41.9s), then rises sharply on heavier agentic workloads like Vals Index (323.2s), Terminal-Bench 2.0 (855.3s), and Finance Agent (5360.3s).
- Overall, Gemma lands at 38.94% on Vals Index and 45.12% on Vals Multimodal Index. The main story here is efficiency for size, not broad frontier dominance.