Release Date: Dec 11, 2025

Developer OpenAI πŸ‡ΊπŸ‡Έ
Context Window 400k
Max Output Tokens 128k
Token Costs (in/out) $1.75/14.00
Weights Private
Input Modalities

Accuracy

71.06 %

Avg. Cost (In/Out)

$ 1.75 / $ 14.00

Latency

16 min 51 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±2.26
13/84

0.0%

Β±1.86
22/83

0.0%

Β±0.92
29/94

0.0%

Β±3.35
19/75

0.0%

Β±0.84
10/139

0.0%

Β±5.07
27/83

0.0%

Β±1.84
17/132

0.0%

Β±8.32
8/62

0.0%

Β±1.03
29/136

0.0%

Β±0.40
54/136

0.0%

Β±0.34
42/132

0.0%

Β±0.82
14/88

0.0%

Β±1.92
37/82
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Reasoning Effort: xhigh

Updates

Dec 11, 2025

GPT 5.2 is the new state of the art on our Vals Index, showing strong performance across domains, especially coding.

Most impressive was the model setting a new state-of-the-art on Vibe Code Bench by an extremely large margin, from 24.6% to 41.31%. It also got first on our IOI, Terminal-Bench, and SWE-bench Verified.

This performance improvement does come with an increased cost across the board - the model is priced at 1.75/1.75 / 14, compared to 1.25/1.25 / 10 for its predecessor. It also tends to see increased token usage, especially on longer running agentic tasks. With the β€œhigh” or the new β€œxhigh” reasoning modes enabled, you may also see very long response times - for particularly tricky questions, it would think for more than 30 minutes.

Overall, the model will be a powerhouse for users seeking strong performance and reliability, particularly on complex reasoning and coding tasks.