May 19, 2026
Google's Gemini 3.5 Flash evaluated across our benchmark suite
-
We evaluated Googleβs new Gemini 3.5 Flash across our benchmark suite. It is the new #1 model on our Finance Agent Benchmark v2, dethroning GPT 5.5 by 6 points.
-
The model places #3 on the Vals Index, scoring 62.05%, and #3 on the Vals Multimodal Index, scoring 62.29%. That is an ~8-point increase over Gemini 3.1 Pro Preview (02/26) on the Vals Index and a ~7-point increase on the Vals Multimodal Index.
-
We also saw a meaningful increase on Vibe Code Bench, our benchmark measuring how well models can create web applications from scratch. Gemini 3.5 Flash ranked #10, scoring 48.68% β a 16-point jump relative to Gemini 3.1 Pro Preview (02/26). The model also did well on other coding benchmarks, ranking #3 on SWE-bench Verified and #3 on Terminal Bench 2.0.
-
On MedCode, our benchmark that assesses whether models can support the medical billing process, Gemini 3.5 Flash ranks #3 overall β just 3.2 points behind the top model, Gemini 3.1 Pro Preview (02/26). It also had a slight 3-point increase on ProofBench.
-
The model still lags behind Gemini 3.1 Pro Preview (02/26) on some of our benchmarks: LegalBench, MortgageTax, and academic benchmarks including MMLU Pro, GPQA.
The model has a 1M-token context window. Evaluations were run using a reasoning effort of βhighβ, a temperature of 1.0, and max output tokens set to 65k.
Congrats to the Google team on the strong release!