Sep 22, 2026
OpenAI's GPT-6 Sol evaluated across our benchmark suite
We evaluated OpenAIโs new GPT-6 Sol across our benchmark suite.
-
GPT-6 Sol places #8 of 65 on the Vals Index (62.57%) at $7.56 per test, just behind GPT-5.6 Sol (63.71%) and about four points behind GPT-6 Astra (66.61%), at well under half of Astraโs $19.09 per test.
-
Top-ten finishes on BioMysteryBench (74.81%, #4 of 16), Code Migration (57.20%, #4 of 68), Terminal-Bench 2.1 (83.15%, #6 of 73), Vibe Code Bench (87.82%, #6 of 103), ProgramBench (2.00% fully resolved, #6 of 51), IOI (82.61%, #7 of 32), EMB (71.53%, #8 of 65) and ProofBench v1.1 (83.00%, #9 of 40).
-
Mid-board or below on SNAP (56.63%, #30 of 40), MedCode (47.07%, #31 of 98), Finance Agent v2 (49.05%, #37 of 68), Legal Research Bench (28.85%, #38 of 68), SAGE (44.79%, #39 of 85), MedScribe (82.03%, #45 of 100), Tax Agent Bench (53.05%, #22 of 26) and Harveyโs Legal Agent Benchmark (1.67%, #41 of 69, where every task requires a fully correct end-to-end answer).
No fallback models were used; refusals and provider policy blocks are counted as failed tasks.
GPT-6 Sol is priced at $2.00 per million input tokens and $10.00 per million output tokens ($4.00 / $15.00 for long-context requests). The model has a 1M-token context window and 128k max output tokens. Evaluations were run with reasoning effort set to โmaxโ.
Congrats to the OpenAI team on the release!