Sep 24, 2025
GPT 5 Codex Evaluated on Coding Benchmarks
We evaluated GPT 5 Codex across Terminal-Bench, SWE-bench Verified, IOI v1, and LCB, finding the following:
- On Terminal-Bench, GPT-5 Codex takes 1st place with 58.8% accuracy, a 10% improvement over the previous #1, GPT-5. It also delivers lower cost and latency compared to the other top three models.
- GPT 5 Codex also gets first place on SWE-bench Verified, narrowly outperforming GPT 5, which ranks 2nd by less than a percentage point.
- GPT 5 Codex is the first model weβve seen receive full credit on a single question on IOI v1, though overall accuracy remains low (9.8%).
- GPT 5 Codex places second on LCB, behind GPT 5 Mini.
GPT 5 Codex is optimized for agentic coding, particularly within OpenAIβs Codex offering. For standardization, we used the same prompts and templates as we used with other models when running our evaluations.