Release Date: Apr 16, 2025

Developer OpenAI πŸ‡ΊπŸ‡Έ
Context Window 200k
Max Output Tokens 100k
Token Costs (in/out) $2.00/8.00
Weights Private
Input Modalities

Accuracy

72.38 %

Avg. Cost (In/Out)

$ 2.00 / $ 8.00

Latency

34.90 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±2.16
23/84

0.0%

Β±1.87
53/83

0.0%

Β±0.93
38/94

0.0%

Β±3.31
43/75

0.0%

Β±0.85
31/139

0.0%

Β±1.86
48/132

0.0%

Β±1.03
40/136

0.0%

Β±0.42
40/136

0.0%

Β±0.34
51/132

0.0%

Β±0.95
39/88
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 100,000
Reasoning Effort: high

Updates

Apr 18, 2025

We just evaluated o3 and o4 Mini on all benchmarks!

  • o3 achieved the #1 overall accuracy ranking on our benchmarks, with exceptional performance on complex reasoning tests like MMMU Pro (#1/22), MMLU Pro (#1/35), GPQA Diamond (#1/35) and proprietary benchmarks like TaxEval (#1/42) and CorpFin (#2/35).

  • o4 Mini achieved the second-highest accuracy across our benchmarks (82.8%), driven by strong performance on public math tests like MGSM (#1/36), MMMU Pro (#2/22), and Math500 (#4/38).

  • Legal benchmark weaknesses: Both models demonstrated significant weaknesses on our proprietary legal benchmarks, with lower ranks on ContractLaw (o3: #34/62, o4 Mini: #14/62) and CaseLaw (o3: #15/55, o4 Mini: #18/55).

  • Cost-effectiveness comparison: With similar performance levels, cost becomes a key differentiator. o4 Mini costs 4.40foroutput,comparedto4.40 for output, compared to 40.00 for o3 β€” a tenfold price difference that makes o4 Mini the more economical choice for many use cases.