Release Date: Jan 31, 2025

Developer OpenAIΒ πŸ‡ΊπŸ‡Έ
Context Window 200k
Max Output Tokens 100k
Token Costs (in/out) $1.10/4.40
Weights Private
Input Modalities

Accuracy

73.33 %

Avg. Cost (In/Out)

$ 1.10 / $ 4.40

Latency

52.50 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.91
96/141

0.0%

Β±2.18
79/135

0.0%

Β±1.15
82/140

0.0%

Β±0.72
112/139

0.0%

Β±0.39
100/135
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 100,000
Reasoning Effort: high

Updates

Feb 3, 2025

We just evaluated OpenAI’s o3-mini model!

  • The model shows a good price-performance trade-off, reaching close to top places on our most recent and proprietary benchmarks like Tax Eval.
  • However, o3-mini seems to struggle with large context windows, performing poorly on the Max Fitting Context task of CorpFin. It tends to lose the question if it is provided at the beginning of a large context window (around 150k tokens and more).

We have also run DeepSeek R1 on our CorpFin benchmark, on which it reaches the top place, beating all other models we have tested.