Apr 7, 2025
Llama 4 Maverick and Llama 4 Scout evaluated on all benchmarks!
We just evaluated Llama 4 Maverick and Llama 4 Scout on all benchmarks!
- Llama 4 Scout achieves an average accuracy of 61.5% with a latency of 7.13 seconds, placing the model at a tie with Mistral Small 3.1 (03/2025) (61.5%) and just behind Cohereโs Command A (63.5%).
- Llama 4 Maverick sits at 67.0% accuracy with a latency of 7.72 seconds ranking just behind Anthropicโs Claude 3.5 Sonnet (69.9%) and DeepSeek V3 (03/24/2025) (74.7%).
- Both models excel on public benchmarks, with Maverick achieving top rankings in MMMU Pro (4/17), MGSM (4/28), GPQA Diamond (5/27), and MMLU Pro (5/27), while Scout delivers strong results in MMMU Pro (10/17), MortgageTax (10/18), and AIME (11/26).
- However, these models show a significant gap between their impressive public benchmark performance and mediocre results on private benchmarks, particularly struggling with TaxEval (Maverick: 28/34, Scout: 32/34), Contract Law (Maverick: 37/54, Scout: 43/54), and MedQA (Maverick: 32/32, Scout: 30/32).