Release Date: Apr 14, 2025

Developer OpenAIΒ πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 33k
Token Costs (in/out) $2.00/8.00
Weights Private
Input Modalities

Accuracy

63.96 %

Avg. Cost (In/Out)

$ 2.00 / $ 8.00

Latency

1 min 25 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.92
37/95

0.0%

Β±0.84
22/141

0.0%

Β±2.40
98/133

0.0%

Β±1.12
109/138

0.0%

Β±0.43
49/137

0.0%

Β±0.39
84/133

0.0%

Β±1.07
61/89
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 32,768
Reasoning Effort: high

Updates

Apr 15, 2025

We just evaluated GPT 4.1, GPT 4.1 Mini, and GPT 4.1 Nano on all benchmarks!

  • GPT 4.1 delivers impressive results with a 75.5% average accuracy across benchmarks.

  • Impressive performance on proprietary benchmarks! GPT 4.1 is now the leader on CorpFin (71.2%), and shows strong performance on CaseLaw (85.8%, 4/53), and MMLU Pro (80.5%, 6/33).

  • GPT 4.1 Nano and GPT 4.1 Mini bring AI to time-sensitive applications with an outstanding latency of only 3.62s and 6.60s respectively while still achieving 59.1% and 75.1% average accuracy.

  • Compact but capable! Despite its size, GPT 4.1 Mini performs admirably on Math500 (88.8%, 10/36) and MGSM (87.9%, 20/34).

  • Size versus performance tradeoff: The smaller models do show lower performance on some complex tasks, with GPT 4.1 Nano ranking near the bottom on MMLU Pro (62.3%, 30/33) and MGSM (69.8%, 32/34).