Jul 9, 2026
OpenAI's GPT-5.6 Sol and Terra evaluated across our benchmark suite
We evaluated OpenAIโs new GPT-5.6 Sol and GPT-5.6 Terra across our benchmark suite.
-
Sol ranks #3 on the Vals Index (63.71%) and the Vals Multimodal Index (72.19%).
-
Agentic coding is the headline: Sol takes #1 on SWE-bench Verified (96.20%) and Terminal-Bench 2.1 (85.77%). It also scores 80.50% (#4) on Vibe Code Bench and takes #1 on ProofBench (77.00%).
-
Sol also leads CyberBench (88.14%), the Excel Modeling Benchmark (72.34%), and Legal Research Bench (48.08%).
-
Terra scores 56.54% on the Vals Index and 65.07% on the Vals Multimodal Index, with strong results on ProofBench (71.00%, #3), Legal Research Bench (40.87%, #4), and Terminal-Bench 2.1 (73.41%, #7).