Nov 10, 2024
Results for the new 3.5 Sonnet (Latest) model
- On Legalbench, itโs now exactly tied with GPT 4o, and beats 4o on CorpFin and CaseLaw
- It usually, but not always, performs a few percentage points better than the previous version - for example, on Legalbench (+1.3%), ContractLaw Overall (+0.5%), and CorpFin (+0.8%).
- There are some instances where it experienced a performance regression - including TaxEval Free Response (-3.2%) and CaseLaw Overall (-0.1%).
- Although itโs competitive with 4o, itโs still not at the level of GPT o1, which still claims the top spots on almost all of our leaderboards.