Release Date: Aug 7, 2025

Developer OpenAIΒ πŸ‡ΊπŸ‡Έ
Context Window 400k
Max Output Tokens 128k
Token Costs (in/out) $1.25/10.00
Weights Private
Input Modalities

Accuracy

66.74 %

Avg. Cost (In/Out)

$ 1.25 / $ 10.00

Latency

5 min 5 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±2.10
16/92

0.0%

Β±1.94
34/94

0.0%

Β±0.88
42/98

0.0%

Β±3.35
40/81

0.0%

Β±0.87
49/145

0.0%

Β±3.41
64/95

0.0%

Β±1.76
46/138

0.0%

Β±0.97
28/143

0.0%

Β±0.38
14/144

0.0%

Β±0.34
41/138

0.0%

Β±0.93
39/93

0.0%

Β±2.07
67/88

0.0%

Β±5.15
39/67
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Reasoning Effort: high

Updates

Aug 7, 2025

We evaluated OpenAI’s newly-released GPT 5 family of models and found that GPT 5 achieves SOTA performance for a fraction of the cost compared to similarly performing models.

Of the three, GPT 5 is the strongest model in the family with SOTA performance on public LegalBench and AIME benchmarks.

On private benchmarks, GPT 5 and GPT 5 Mini achieve top 10 performance on all but CaseLaw. Most notably, GPT 5 Mini is the new SOTA model on TaxEval at a substantially lower cost.

On public benchmarks, GPT 5 places top 5 and GPT 5 Mini places top 10 on nearly everything. Further, we found the two models have complementary strengths - GPT 5 is SOTA on LegalBench and AIME, while GPT 5 Mini is SOTA on LiveCodeBench.

Lastly, GPT 5 Nano achieves middle of the pack performance across the board. It narrowly places in the top 10 on AIME, compared to GPT 5 and GPT 5 Mini which top the charts.