Archived Benchmark

Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases. Previous results are preserved here for posterity.

Proprietary

CorpFin

Updated 8/12/2026

A private benchmark evaluating understanding of long-context credit agreements

As of August 12, 2026, Claude Opus 5 ranks first on CorpFin with 73.19%, followed by Claude Fable 5 (71.83%) and Kimi K3 (71.56%).

CorpFin v2Long credit agreement comprehension
ACCURACY

CorpFin leaderboard

Rank Model Accuracy Cost In / Out Latency
1 Claude Opus 5 73.19% $5 / $25 28.57s
2 Claude Fable 5 71.83% $10 / $50 52.99s
3 Kimi K3 71.56% $3 / $15 37.99s
4 Muse Spark 1.1 71.29% $1.25 / $4.25 27.82s
5 Muse Spark 1.2 70.94% $1.25 / $4.25 30.05s
6 Inkling Small 69.62% $0.3 / $1.2 45.32s
7 Inkling 68.57% $1 / $4.05 29.20s
8 Grok 4.3 68.53% $1.25 / $2.5 36.44s
9 GPT 5.5 68.42% $5 / $30 23.30s
10 Kimi K2.5 68.26% $0.6 / $3 80.78s
11 MiniMax-M3 68.10% $0.6 / $2.4 24.46s
12 Qwen 3 Max Thinking 68.03% $1.2 / $6 2m01s
13 Grok 4.5 67.41% $2 / $6 91.88s
14 Claude Opus 4.6 (Thinking) 67.02% $5 / $25 20.58s
15 Claude Sonnet 5 66.98% $2 / $10 18.63s
16 Grok 4 Fast (Reasoning) 66.90% $0.2 / $0.5 11.84s
17 Kimi K2.6 66.74% $0.95 / $4 79.59s
18 Claude Opus 4.8 66.71% $5 / $25 20.23s
19 Qwen 3.6 Max Preview 66.47% $1.3 / $7.8 47.07s
20 Gemini 3 Flash (12/25) 66.43% $0.5 / $3 11.28s
21 Grok 4.6 66.16% $2 / $6 24.26s
22 GLM 5.2 66.12% $1.4 / $4.4 24.59s
23 Claude Opus 4.7 66.08% $5 / $25 17.75s
24 Grok 4 66.05% $3 / $15 29.73s
25 Grok 4.1 Fast (Reasoning) 65.97% $0.2 / $0.5 28.41s
26 GPT 5.2 65.89% $1.75 / $14 26.04s
27 Qwen 3.8 Max 65.85% $2 / $6 46.65s
28 Nemotron 3 Ultra 65.46% N/A 15.02s
29 DeepSeek V4 Pro 0813 65.42% $1.32 / $3.96 32.85s
30 GPT-5.6 Terra 65.31% $2 / $12 7.72s
31 Claude Sonnet 4.6 65.31% $3 / $15 18.48s
32 Qwen 3.5 Plus 65.31% $0.4 / $2.4 2m36s
33 GPT 5.4 (xhigh) 65.27% $2.5 / $15 82.34s
34 Muse Spark 65.11% N/A 40.06s
35 Claude Opus 4.5 (Thinking) 65.07% $5 / $25 17.73s
36 Gemini 3.5 Flash 64.69% $1.5 / $9 8.70s
37 Gemini 3.1 Pro Preview (02/26) 64.49% $2 / $12 25.39s
38 GLM 5.1 64.45% $1 / $3.2 68.17s
39 GPT-5.6 Sol 64.38% $4 / $20 15.94s
40 GPT-5.6 Luna 64.22% $0.2 / $1.2 16.97s
41 GPT 5.1 63.83% $1.25 / $10 35.29s
42 Qwen 3.7 Max 63.71% $2.5 / $7.5 19.93s
43 Grok 4.20 (Reasoning) 63.67% $2 / $6 9.67s
44 Gemini 3 Pro (11/25) 63.67% $2 / $12 20.06s
45 Qwen 3.5 Flash 63.56% $0.1 / $0.4 35.20s
46 Gemini 3.6 Flash 63.33% $1.5 / $7.5 6.33s
47 GPT 4.1 63.05% $2 / $8 10.14s
48 GLM 5 62.90% $1 / $3.2 68.58s
49 Qwen 3.6 27B 62.31% $0.6 / $3.6 62.76s
50 Claude Sonnet 4.5 (Thinking) 61.97% $3 / $15 20.54s
51 Ling 3.0 Flash 61.93% $0.075 / $0.22 12.45s
52 Qwen 3.6 Plus 61.93% $0.5 / $3 35.84s
53 DeepSeek V4 Flash 0731 61.85% $0.44 / $1.32 16.48s
54 MiMo V2.5 Pro 61.42% $0.435 / $0.87 18.90s
55 DeepSeek V4 61.38% $1.32 / $3.96 79.33s
56 Claude Opus 4.5 (Nonthinking) 61.30% $5 / $25 9.40s
57 Claude Sonnet 4 (Thinking) 61.23% $3 / $15 2m49s
58 GPT 5.4 Nano 61.19% $0.2 / $1.25 9.40s
59 MiniMax-M2.7 61.19% $0.3 / $1.2 10.97s
60 Grok 3 Mini Reasoning 61.11% $0.3 / $0.5 14.39s
61 GPT 5 61.07% $1.25 / $10 2m41s
62 Mistral Large 3 61.03% $0.5 / $1.5 48.84s
63 GLM 4.5 60.96% $0.6 / $2.2 95.24s
64 GPT 5.4 Mini 60.92% $0.75 / $4.5 52.30s
65 Claude Sonnet 4.5 (Nonthinking) 60.80% $3 / $15 13.38s
66 Gemini 2.5 Pro 60.80% $1.25 / $10 27.43s
67 Gemini 3.5 Flash Lite 60.68% $0.3 / $2.5 3.84s
68 Claude Haiku 4.5 (Thinking) 60.61% $1 / $5 14.24s
69 Kimi K2 Thinking 60.57% $0.6 / $2.5 93.30s
70 Claude 3.7 Sonnet (Thinking) 60.41% $3 / $15 22.16s
71 Claude Haiku 4.5 (Nonthinking) 60.30% $1 / $5 7.95s
72 GPT 5 Mini 60.18% $0.25 / $2 73.79s
73 MiMo V2.5 59.91% $0.14 / $0.28 17.63s
74 Gemini 2.5 Pro Exp 59.83% $1.25 / $10 16.30s
75 Gemini 2.5 Flash Preview (9/25) (Thinking) 59.75% $0.3 / $2.5 51.76s
76 o3 59.71% $2 / $8 19.15s
77 Grok 3 59.71% $3 / $15 24.35s
78 MiniMax-M2.5 59.60% $0.3 / $1.2 16.60s
79 Grok 3 Mini Reasoning 59.48% $0.3 / $0.5 10.52s
80 Gemini 3.1 Flash Lite Preview 59.36% $0.25 / $1.5 4.03s
81 o4 Mini 58.97% $1.1 / $4.4 17.53s
82 Gemini 2.5 Flash Preview (9/25) (Nonthinking) 58.97% $0.3 / $2.5 51.90s
83 MiniMax-M2.1 58.90% $0.3 / $1.2 21.97s
84 Mistral Medium 3.5 58.78% $1.5 / $7.5 43.70s
85 Grok 4 Fast (Non-Reasoning) 58.39% $0.2 / $0.5 9.37s
86 GPT OSS 120B 58.24% $0.15 / $0.6 42.72s
87 Laguna M.1 58.16% N/A 3m32s
88 GPT 4.1 Mini 57.93% $0.4 / $1.6 6.44s
89 Gemini 2.5 Flash Lite (9/25) (Thinking) 57.58% $0.1 / $0.4 54.23s
90 GLM 4.6 56.84% $0.6 / $2.2 78.06s
91 Laguna XS.2 56.33% N/A 3m08s
92 Gemini 2.5 Flash Lite (9/25) (Nonthinking) 56.29% $0.1 / $0.4 54.75s
93 Qwen 3 Max 55.94% $1.2 / $6 55.67s
94 DeepSeek V3 (03/24/2025) 54.74% $0.9 / $0.9 36.73s
95 Claude Sonnet 4 (Nonthinking) 54.70% $3 / $15 3m06s
96 Trinity Large Thinking 54.66% $0.25 / $0.9 12.48s
97 Nemotron 3.5 Lightning 54.23% $0.05 / $0.2 38.79s
98 Gemini 2.5 Flash Preview 4/17 (Nonthinking) 54.16% $0.3 / $2.5 8.99s
99 DeepSeek R1 54.12% $1.35 / $5.4 46.77s
100 Claude 3.5 Sonnet Latest 53.61% $3 / $15 0.05s
101 GPT OSS 20B 53.15% $0.07 / $0.3 33.97s
102 Qwen 3 Max Preview 52.95% $1.2 / $6 3m47s
103 Grok 4.1 Fast Non-Reasoning 52.49% $0.2 / $0.5 13.46s
104 DeepSeek V3 52.49% $0.9 / $0.9 28.32s
105 DeepSeek V3.1 51.48% $0.56 / $1.68 63.00s
106 Grok 2 51.13% $2 / $10 87.13s
107 DeepSeek V3.2 (Thinking) 50.97% $0.56 / $1.68 99.48s
108 Claude 3.5 Haiku Latest 50.82% $0.8 / $4 0.04s
109 Mistral Medium 3.1 (05/2025) 50.74% $0.4 / $2 30.35s
110 Kimi K2 Instruct 50.39% $1 / $3 40.89s
111 Llama 4 Maverick 49.73% $0.22 / $0.88 4.72s
112 DeepSeek V3.2 (Nonthinking) 47.94% $0.56 / $1.68 2m14s
113 Magistral Medium 1.2 (09/2025) 47.40% $2 / $5 55.27s
114 Llama 4 Scout 46.78% $0.18 / $0.59 9.70s
115 GLM 4.7 46.39% $0.6 / $2.2 66.10s
116 Command A 45.96% $2.5 / $10 14.15s
117 GPT 4o (2024-11-20) 45.92% $2.5 / $10 6.46s
118 GPT 4o Mini 45.45% $0.15 / $0.6 0.04s
119 o3 Mini 45.30% $1.1 / $4.4 31.05s
120 Mistral Small 3.1 (03/2025) 44.17% $0.075 / $0.3 11.29s
121 Magistral Small 1.2 (09/2025) 44.02% $0.5 / $1.5 14.05s
122 Gemini 2.0 Pro Exp 43.44% $1.25 / $5 18.88s
123 GPT 4.1 Nano 42.08% $0.1 / $0.4 4.95s
124 Jamba 1.6 Large 41.53% $2 / $8 27.38s
125 Gemini 1.5 Pro (002) 40.52% $1.25 / $5 38.05s
126 GPT 4o (2024-08-06) 39.43% $2.5 / $10 0.05s
127 Jamba 1.5 Large 39.43% $2 / $8 10.77s
128 Llama 3.1 Instruct Turbo (70B) 38.85% $0.88 / $0.88 0.06s
129 Gemini 1.5 Flash (002) 38.19% $0.075 / $0.3 28.41s
130 Jamba 1.6 Mini 38.03% $0.2 / $0.4 4.25s
131 Llama 3.1 Instruct Turbo (8B) 37.80% $0.18 / $0.18 0.04s
132 Jamba 1.5 Mini 33.88% $0.2 / $0.4 2.31s
133 Gemini 2.0 Flash (001) 33.72% $0.1 / $0.4 30.04s
134 Gemini 1.5 Flash (001) 28.63% $0.075 / $0.3 1.16s

Key Takeaways

  • The top of the leaderboard is a near dead heat — the top six models fall within about 3.5 points (led by Claude Opus 5 at 73.19%).
  • Long context is the real differentiator: several models stumble on the Max Fitting Context task, losing track of the question when it appears at the start of a very large context window (around 150k tokens and up).

Dataset and Context

In both the finance and legal industries, it is a common task to ask a specific question or understand a piece of information from a very long document. A common document type is a credit agreement - contracts, often over 200 pages, used when large corporations receive a line of credit from a banking entity (see this example credit agreement filed with the SEC).

For this dataset, we worked with a team of experts, including financial analysts, legal professionals, and academics, to create a set of questions and answers about these credit agreements. The dataset is divided into:

  • Public Validation: 20 questions from 1 document, accessible to anyone on request (email contact@vals.ai for access).
  • Private Validation: 340 questions from 17 documents, available for purchase to evaluate and improve your own models.
  • Test: 858 questions from 43 documents. This is the privately held set for Vals AI benchmarking and is never shared.

The dataset is also organized into three distinct tasks, which all use the entire set of questions, but pass a different subset of information into each model’s context window. The three tasks are:

  1. Exact Pages: This variation provides only the necessary pages required to answer each question, typically resulting in a small context of only a few pages.
  2. Shared Max Context: This variation takes a subset of pages (~80) that is guaranteed to fit in the context window of all models, and also contain the answer. The subselection doesn’t necessarily start at the first page, which can make document structure comprehension challenging for models.
  3. Max Fitting Context: This variation includes the largest possible chunk of the document, starting from the first page, that fits within the model’s context window. This means longer-context models have more information than models with shorter context windows.

The types of questions written by the experts include:

  • Basic extraction of terms and numbers: Some examples include “Who is the borrower’s legal counsel?” or “What currencies are the USD 325mm RCF first lien available in?”
  • Summarization and interpretation questions: Some examples include, “Is there an erroneous payment provision?” and “How will the loan proceeds be used?”
  • Numeric reasoning or calculation-based questions: Some examples include “How much initial debt capacity is available to the Borrower on day one?”
  • Questions involving referring to multiple sections of the provided text, especially to previous definitions: For instance, one example is “What is the minimum required amount of Available Liquidity that the company must maintain?” The model must refer to a previous definition of Available Liquidity.
  • Giving opinions of terms based on market standards: For instance, “Are there any unusual terms used to define or adjust EBITDA?” These questions require the models to make a judgment call, rather than just make a statement of fact.
  • Questions making use of industry jargon: There are several terms like “baskets”, which have a commonly understood meaning in the industry, but are almost never explicitly used in the agreement itself. An example is “Does the contract contain a Chewy Blocker?” (a type of clause meant to prevent a subsidiary from being released from its debt obligations).

Results

Although the larger and more powerful flagship models generally perform the best, the gap between the smaller models is not as big as in other benchmarks. When the latency and price are taken into account, it means that sometimes the smaller models provide a better accuracy for a given cost.



Model Output Examples

The questions that require synthesizing content and calculating values across multiple sections in the document are some of the hardest. The fact that the text is parsed from PDFs also poses difficulty for the models: they need to handle special characters, tables, etc.

Below is an example of a number retrieval question, on the Max Fitting Context task. This question can be difficult for models because in a long financial document, there can be several similar tables that relate to the question, but only one contains the correct information.

In this case, the expected answer is 3.25 to 1.00.

Question

What is the Total Net Leverage Ratio limit for unlimited RPs and investments?

3.25 to 1.00

CORRECT


Additional Notes

Context Window

Our testing across different context lengths reveals that the context window is an extremely significant factor in the model’s performance. You can find the context window for a given model on the provider’s website.

Evaluation Methodology

All evaluations were conducted using Claude 4.5 Sonnet as the judge, with temperature set to 0.

Retrieval-Augmented Generation (RAG)

As an alternative to including the entire document in the context window, RAG techniques are extremely common. This method breaks up long documents and databases into “chunks” which are first retrieved, and then passed to the model as context. We do not test RAG systems in this benchmark.

Evaluation Model Update

The LLM-as-judge for this benchmark was upgraded from Sonnet 3.5 (which is now deprecated and inaccessible) to Sonnet 4.5 on 11/17/25.