Sep 30, 2026
Google's Gemini 4 Argon evaluated across our benchmark suite
We evaluated Googleโs new Gemini 4 Argon across our benchmark suite.
-
Gemini 4 Argon places #1 of 41 on the Vals Index (68.90%) at $15.68 per test, ahead of Claude Sonnet 5.5 (67.04%, $21.34 per test), Claude Opus 5.5 (66.97%, $32.14) and Claude Fable 5.1 (65.83%, $28.71).
-
It takes #1 on Finance Agent v2 (65.40% of 73, ahead of Gemini 3.8 Flash at 61.44%) and ties GPT-6 Astra for #1 on IOI with a perfect 100.00%.
-
It sits at #2 on Vibe Code Bench (91.91% of 106, behind Sonnet 5.5 at 92.39%), Code Migration (68.17% of 71, behind Sonnet 5.5 at 69.83%), Tax Agent Bench (76.23% of 64, behind Fable 5.1 at 77.64%) and CyberBench (77.86% of 43, 0.12 points behind GPT-6 Sol), and at #3 on LegalBench (88.30% of 149), MedCode (58.80% of 104), Terminal-Bench Science (44.29% of 34) and SRE Bench (44.28% of 16).
-
Further top-ten finishes on Legal Research Bench (54.81%, #4 of 72), EMB (75.24%, #4 of 69), the Vals RSI Index (30.55%, #4 of 23), Terminal-Bench 4.0 (57.58%, #5 of 42), Harveyโs Legal Agent Benchmark (19.58%, #5 of 73), SAGE (53.65%, #5 of 90), Public Benefits Bench (69.76%, #5 of 45), ProgramBench (2.50% fully resolved, #5 of 54), ProofBench v1.1 (99.00%, #5 of 46, one point behind a four-way tie at 100.00%), BioMysteryBench (76.30%, #6 of 21) and MysteryMechanism (45.50%, #6 of 21).
-
Weaker on MedScribe (87.43%, #15 of 106) and CUA-bench (4.83%, #7 of 8).
-
Cost per test on the Vals Index sits below the three Claude models that follow it ($15.68 versus $21.34 to $32.14), but climbs on long agentic tasks: $193.78 per test on CUA-bench, $57.82 on Code Migration, $44.90 on Terminal-Bench Science and $30.45 on SRE Bench.
Gemini 4 Argon is priced at $4.00 per million input tokens and $20.00 per million output tokens. The model has a 1M-token context window and 262k max output tokens. Evaluations were run with reasoning effort set to โhighโ and temperature set to 1.0.
Congrats to the Google team on the release!