Release Date: Mar 18, 2025

Developer NVIDIA πŸ‡ΊπŸ‡Έ
Context Window 131k
Max Output Tokens 33k
Token Costs (in/out) $0.00/0.00
Weights Open
Input Modalities

Accuracy

63.51 %

Avg. Cost (In/Out)

N/A

Latency

1 min 15 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.94
115/139

0.0%

Β±2.45
102/132

0.0%

Β±1.17
102/136

0.0%

Β±0.44
118/132
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : NVIDIA
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 32,768

Updates

Jul 29, 2025

We evaluated Llama 3.3 Nemotron Super (Nonthinking) and Llama 3.3 Nemotron Super (Thinking) and found the Thinking variant substantially outperforms the Nonthinking variant, with the significant exception of our proprietary Contract Law benchmark.