Release Date: Mar 5, 2026

Developer OpenAIΒ πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $2.50/15.00
Weights Private
Input Modalities

Accuracy

61.38 %

Avg. Cost (In/Out)

$ 2.50 / $ 15.00

Latency

13 min 45 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±4.11
22/59

0.0%

Β±5.65
10/25

0.0%

Β±2.15
44/92

0.0%

Β±3.32
56/94

0.0%

Β±0.92
17/98

0.0%

Β±3.12
41/81

0.0%

Β±0.87
44/145

0.0%

Β±4.84
26/95

0.0%

Β±1.91
21/138

0.0%

Β±1.04
41/143

0.0%

Β±0.41
13/144

0.0%

Β±0.42
27/138

0.0%

Β±0.80
15/93

0.0%

Β±0.00
23/45

0.0%

Β±4.72
19/35

0.0%

Β±1.85
30/88

0.0%

Β±5.25
12/67
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Reasoning Effort: xhigh

Updates

Mar 5, 2026

We evaluated GPT 5.4 (xhigh) across our full benchmark suite.

Evaluations were run via the official OpenAI API using β€œxhigh” reasoning, except on Terminal-Bench 2.0, where we used β€œhigh”.