Proprietary

Finance Agent

Updated 6/4/2026

Evaluating agents on core financial analyst tasks

As of June 4, 2026, Claude Opus 4.7 ranks first on Finance Agent v1.1 with 64.37%, followed by Claude Sonnet 4.6 (63.33%) and Muse Spark (60.59%).

Finance Agent v1.1
ACCURACY

Finance Agent leaderboard

Rank Model Accuracy Cost / Test Input / Output Cost Duration
1 Claude Opus 4.7 64.37% $0.80 $5 / $25 4m31s
2 Claude Sonnet 4.6 63.33% $1.44 $3 / $15 5m49s
3 Muse Spark 60.59% N/A N/A 6m58s
4 DeepSeek V4 60.39% $0.59 $1.32 / $3.96 9m48s
5 Claude Opus 4.6 (Thinking) 60.05% $1.11 $5 / $25 4m50s
6 GPT 5.5 59.96% $1.33 $5 / $30 14m15s
7 Gemini 3.1 Pro Preview (02/26) 59.72% $0.87 $2 / $12 4m26s
8 Claude Opus 4.5 (Thinking) 58.81% $1.50 $5 / $25 3m02s
9 GPT 5.2 58.53% $0.98 $1.75 / $14 9m45s
10 GLM 5.1 57.66% $0.29 $1 / $3.2 8m22s
11 GPT 5.4 (xhigh) 57.15% $1.41 $2.5 / $15 10m57s
12 Kimi K2.6 57.06% $0.49 $0.95 / $4 25m06s
13 GPT 5.1 55.31% $0.47 $1.25 / $10 9m38s
14 Gemini 3 Pro (11/25) 55.15% $0.56 $2 / $12 3m04s
15 Qwen 3.6 Plus 54.63% $0.15 $0.5 / $3 5m28s
16 Claude Sonnet 4.5 (Thinking) 54.50% $1.10 $3 / $15 3m22s
17 Qwen 3.5 Plus 54.48% $0.24 $0.4 / $2.4 6m00s
18 Grok 4.3 53.81% $0.44 $1.25 / $2.5 13m16s
19 Grok 4 53.51% $1.07 $3 / $15 5m21s
20 GPT 5.4 Mini 53.41% $0.49 $0.75 / $4.5 14m36s
21 GLM 5 53.18% $0.50 $1 / $3.2 9m24s
22 Qwen 3.6 Max Preview 52.78% $0.59 $1.3 / $7.8 30m39s
23 Grok 4.1 Fast (Reasoning) 52.45% $0.06 $0.2 / $0.5 90.58s
24 Grok 4.20 (Reasoning) 52.30% $0.36 $2 / $6 2m07s
25 GPT 5 52.15% $0.59 $1.25 / $10 15m27s
26 GPT 5 Mini 51.93% $0.14 $0.25 / $2 10m44s
27 Gemma 4 31B IT 50.79% Free $0 / $0 1h29m
28 Kimi K2.5 50.62% $0.18 $0.6 / $3 4m29s
29 MiniMax-M2.7 48.40% $0.16 $0.3 / $1.2 21m33s
30 GPT 5.4 Nano 47.80% $0.01 $0.2 / $1.25 4m44s
31 Gemini 3 Flash (12/25) 47.60% $0.37 $0.5 / $3 14m36s
32 Claude Haiku 4.5 (Thinking) 46.93% $0.38 $1 / $5 117.94s
33 Gemini 3.1 Flash Lite Preview 46.12% $0.07 $0.25 / $1.5 102.44s
34 Mistral Medium 3.5 46.11% $1.15 $1.5 / $7.5 9m22s
35 Grok 4 Fast (Reasoning) 46.08% $0.05 $0.2 / $0.5 58.80s
36 GLM 4.7 45.98% $0.24 $0.6 / $2.2 10m05s
37 Qwen 3.5 Flash 45.64% $0.05 $0.1 / $0.4 4m09s
38 Grok 4.1 Fast Non-Reasoning 44.36% $0.09 $0.2 / $0.5 93.21s
39 Qwen 3 Max 44.30% $0.80 $1.2 / $6 5m15s
40 Gemini 2.5 Pro 41.59% $0.34 $1.25 / $10 30m40s
41 MiniMax-M2.5 38.58% $0.16 $0.3 / $1.2 22m35s
42 Kimi K2 Thinking 36.65% $0.25 $0.6 / $2.5 11m07s
43 GLM 4.6 36.48% $0.19 $0.6 / $2.2 12m16s
44 MiniMax-M2.1 33.35% $0.11 $0.3 / $1.2 21m41s
45 GPT OSS 120B 21.54% $0.06 $0.15 / $0.6 3m05s
46 Mistral Large 3 18.05% $0.09 $0.5 / $1.5 91.15s
47 GPT 4o (2024-08-06) 8.06% $0.16 $2.5 / $10 31.14s
48 Command A 4.23% $2.34 $2.5 / $10 9m33s
49 DeepSeek V3.2 (Thinking) 2.35% $0.21 $0.56 / $1.68 28m41s
50 Jamba 1.7 Large 0.37% $5.04 $2 / $8 38m51s
51 DeepSeek V3.2 (Nonthinking) 0.00% $0.04 $0.56 / $1.68 100.62s

Key Takeaways

  • Claude Opus 4.7 is the current top performer on Finance Agent, scoring 64.37% accuracy. Claude Sonnet 4.6 follows with 63.33%, Muse Spark with 60.59%, DeepSeek V4 with 60.39%, and Claude Opus 4.6 (Thinking) with 60.05%.
  • In the last six months, we’ve seen significant improvement on the benchmark - the capability of LLMs to take on financial tasks is dramatically increasing.
  • Models on average performed best in the simple quantitative and qualitative retrieval tasks. These tasks are easy but time-intensive for finance analysts.

Background

The frontier of applied AI is agents – systems that independently direct their own processes to maintain control over how they accomplish tasks on behalf of users [1] [2]. As such, foundation model labs have invested heavily in developing agents that can handle complex tasks [3], making them prime candidates for delivering significant ROI in specialized industries [1] [4].

Finance is one of the most lucrative applications of agents [5], where AI has the potential to drive significant efficiency gains by performing tasks that mirror those of an entry-level financial analyst [6]. Yet, there lacks a framework for evaluating these agents [7], highlighting the need for standardized evaluation methods to measure agentic ability on domain-specific tasks.

We have created a benchmark that tests the ability of agents to perform tasks expected of an entry-level financial analyst. In collaboration with Stanford researchers, a Global Systemically Important Bank, and industry experts, we created a dataset of 537 questions that evaluate skills such as simple retrieval, market research, and projections.


Results

Finance Agent v1.1

The Pareto cost-curve above illustrates the relationship between model accuracy and cost per session across different models. As model performance improves, these cost considerations will likely be more important.

Some challenging queries cost more than $5 for o3 or Claude Opus 4.1 (Thinking). Often, performance can be improved by allowing the model more tool calls and more tokens used, but this comes at a cost.


Tool Calls Analysis
TURNSTOOL CALLSERRORS

The tool call analysis reveals how many steps the models take to produce an answer. Interestingly, the top performers on this benchmark also register relatively high numbers of tool calls. This suggests that their comparatively strong performance on financial analyst tasks is partly due to leveraging the tools provided. It also reveals certain top models are making multiple tool calls per turn.

Tool Usage
6/50 models

The spider graph detailing the number of tool calls performed by each model further supports the notion that the number of tool calls generally leads to higher accuracy. Again, we see several of the top models excelling by making many tool calls, especially for retrieving information, parse html_page, and Google web-search, which leads to top benchmark performance.

Model Output Examples

Question

Calculate the % Change in the Total Number of Common Stock Shares Repurchased by Netflix (NASDAQ: NFLX) in Q4 2024 vs Q4 2023.

## % Change in Total Common Stock Shares Repurchased by Netflix: Q4 2024 vs. Q4 2023

### Data Gathered from Netflix's 10-K Filings (Issuer Purchases of Equity Securities)

**Q4 2024 Share Repurchases** (from Netflix's 10-K for fiscal year 2024, filed January 27, 2025):
| Month | Shares Repurchased | Avg. Price/Share |
|---|---|---|
| October 2024 | 519,883 | $724.15 |
| November 2024 | 457,732 | $792.49 |
| December 2024 | 188,212 | $913.13 |
| **Total Q4 2024** | **1,165,827** | |

**Q4 2023 Share Repurchases** (from Netflix's 10-K for fiscal year 2023, filed January 26, 2024):
| Month | Shares Repurchased | Avg. Price/Share |
|---|---|---|
| October 2023 | 287,360 | $404.62 |
| November 2023 | 2,708,477 | $447.03 |
| December 2023 | 2,481,771 | $472.63 |
| **Total Q4 2023** | **5,477,608** | |

### % Change Calculation

% Change = (Q4 2024 − Q4 2023) / Q4 2023 × 100
= (1,165,827 − 5,477,608) / 5,477,608 × 100
= −4,311,781 / 5,477,608 × 100
= **−78.72%**

### Answer
Netflix's total number of common stock shares repurchased in **Q4 2024 decreased by approximately 78.72%** compared to Q4 2023 (from 5,477,608 shares down to 1,165,827 shares). The decline was driven largely by significantly higher share prices in Q4 2024 (averaging $700–$900+ per share vs. $400–$475 per share in Q4 2023), resulting in fewer shares being repurchased despite continued buyback activity.

{
"sources": [
{
"url": "https://www.sec.gov/Archives/edgar/data/1065280/000106528025000044/nflx-20241231.htm",
"name": "Netflix, Inc. Form 10-K for Fiscal Year Ended December 31, 2024 (Issuer Purchases of Equity Securities)"
},
{
"url": "https://www.sec.gov/Archives/edgar/data/1065280/000106528024000030/nflx-20231231.htm",
"name": "Netflix, Inc. Form 10-K for Fiscal Year Ended December 31, 2023 (Issuer Purchases of Equity Securities)"
}
]
}

1

The right answer we were looking for was “2024 - 1,165,827. 2023 - 5,477,608. The number of shares repurchased by the company decreased by 79% in 2024”. To answer this question, the finance expert did the following:

  • Go to Edgar
  • Search Netflix Filings by Ticker
  • Identify the 10-K for 2024 and the 10-K for 2023
  • Search for Company Purchases of Equity Securities
  • Calculate the Total Number of Shares Purchased from Oct 1, 2024 - Dec 31, 2024
  • Calculate the Total Number of Shares Purchased from Oct 1, 2023 - Dec 31, 2023
  • Calculate the Percent Change in Total Number of Shares Purchased from 2023 to 2024

Trajectory Comparison

The visualizations below illustrate the sequence and types of tool calls made by two models for the example question above. Each node represents a step in the model’s reasoning process, with colored nodes indicating specific tool calls. These trajectories reveal how models approach research tasks, with more thorough tool usage often correlating with better performance.

Claude Sonnet 4.5 (Thinking) trajectory:

Claude 4.5 Sonnet trajectory

We see that Claude Sonnet 4.5 (Thinking)‘s trajectory follows what we’d expect intuitively— an initial edgar_search (corresponding to going to EDGAR and searching for Netflix filings by ticker), followed by parse_html_page, and finally retrieve_information.

Gemini 2.5 Pro Preview trajectory:

Gemini 2.5 Pro trajectory

In the case of Gemini 2.5 Pro Preview, it follows roughly the same pattern as Claude Sonnet 4.5 (Thinking), except the model is also able to recover from a failed tool call!


Methodology

The finance industry comprises a wide array of tasks, but through consultation with experts at banks and hedge funds, we identified one core task shared across nearly all financial analyst workflows: performing research on the SEC filings of public companies. This task —while time-consuming— is foundational to activities such as equity research, credit analysis, and investment due diligence. We collaborated with industry experts to define a question taxonomy, write and review 537 benchmark questions.

The AI agents were evaluated in an environment where they had access to tools sufficient to produce an accurate response. This included an EDGAR search interface via the SEC_API, Google search, a document parser (ParseHTML) for loading and chunking large filings, and a retrieval tool (RetrieveInformation) that enabled targeted questioning over extracted text. The human experts did not make use of any additional tools when writing and answering their questions. See the full harness in the Finance Agent GitHub repository.

Our primary evaluation metric was final answer accuracy (see the GAIA benchmark). We also recorded latency, tool utilization patterns, and associated computational cost to provide a fuller picture of agent efficiency and practical viability. Together, these components form a rigorous and domain-specific evaluation framework for agentic performance in finance, advancing the field’s ability to measure and rely on AI in high-stakes settings.

Page 1 of Edgar Research

The code behind this harness is open source. Dive in and explore it yourself on this repo!

Dataset

The dataset is divided into three parts: Public Validation (50 open-source samples), Private Validation (150 samples available for license), and Test (337 samples).

  • The Public Validation set is fully open.

  • The Private Validation set is available for license. Interested parties are encouraged to contact us directly for access.

  • The Test set will remain private permanently. All results reported in this page are based solely on the Test set to prevent potential future overfitting.

The dataset splits were sampled to preserve the distribution of question types and performance characteristics. We observed a strong correlation in performance across the validation sets and the Test set, supporting the reliability of these splits.


Question Taxonomy

Quantitative Retrieval (easy)

Direct extraction of numerical information from one or more documents without any post-retrieval calculation or manipulation.

What was the quarterly revenue of Salesforce (NYSE:CRM) for the quarter ended December 31, 2024?

Qualitative Retrieval (easy)

Direct quotation or summarization of non-numerical information from one or more documents.

Describe the product offerings and business model of Microsoft (NASDAQ:MSFT)?

Numerical Reasoning (easy)

Calculations or aggregation of key numbers to produce an answer.

What is % of revenue derived from AWS in each year and the 3 year CAGR from 2021-2024 of Amazon?

Complex Retrieval (medium)

Numerical or non-numerical retrieval or content summarization requiring synthesis of information from multiple documents.

Please briefly summarize the most recent capital raise conducted by Viking Therapeutics (NASDAQ:VKTX).

Adjustments (medium)

Quantitative and qualitative analysis of reporting context bridging GAAP and Non-GAAP Financial Metrics.

What is Lemonade Insurance’s Adjusted EBITDA for the year ended December 31, 2024?

Beat or Miss (medium)

Comparison of forward management guidance versus actuals, synthesized by reconciling sequential quarterly reporting documents.

How did Lam Research’s revenue compare to management projections (at midpoint) on a quarterly basis in 2024? Format as % BEAT or MISS. Use guidance provided on a quarterly basis.

Analyze patterns within a single company’s reporting structure or calculate and contextualize evolving performance, key metrics or business composition.

Which Geographic Region has Airbnb (NASDAQ: ABNB) experienced the most revenue growth from 2022 to 2024?

Financial Modeling (hard)

Complex numerical reasoning calculations which require additional financial expertise to define and evaluate.

How much M&A firepower does Amazon have as of FY2024 end including balance sheet cash, non-restricted cash and other short term investments, and up to 2x GAAP EBITDA leverage? Round to nearest billion

Market Analysis (hard)

Advanced analysis of one or more companies using various documents, requiring normalization of comparison metrics, or complex reasoning and usage of causality to contextualize drivers of business changes or competition dynamics.

Compare the quarterly revenue growth of FAANG companies between 2022-2024.


Finance Agent v1.1

At the beginning of 2026, we performed a refresh of the data, the harness, and the evaluation methodology to ensure the benchmark remains at an extremely high standard. The primary changes are as follows:

Data

  • AfterQuery ran quality control on the benchmark data using finance experts from leading global investment banks, private equity firms, and hedge funds, including Goldman Sachs, Silver Lake, and Citadel.
  • We switched many of the questions from relative dates (e.g., “current year”, “last three quarters”) to absolute dates (“April 2025”, “Q1–Q4 2024”). This is in addition to the pre-existing instructions in 1.0, which instruct the model to “answer each question as if today was 4/7/25.”

Harness

  • The search tool was switched from Serp API to Tavily.
  • Model submission is now via a submit tool call rather than extracted from the message history.
  • The prompt template was updated to better instruct the model on including supporting reasoning and evidence in its final answer.
  • The CIK parameter in the SEC search is now optional.
  • Additional instructions on how to use the data storage were added to the prompt template.
  • The system prompt was also modified to instruct the model to give answers to two decimal places and not round any intermediate calculations.

Evaluation

  • The evaluator model was upgraded to GPT-5.2.
  • The LLM-as-judge now uses the mode of three evaluations to reduce variance.
  • Clearer guidelines on dealing with rounding were added to the LLM-as-Judge prompt.

All models were re-run from scratch using the FAB v1.1 harness and data.


Acknowledgements

Thanks to the following people for their support: Shirley Wu, Alfston Thomas, Andrew Schettino, Kathy Ye, Kyle Jung, Matthew Friday, Michael Xia, and Nicholas Crawley-Brown.

For v1.1, thank you to Spencer Mateega, Sam Jacob, and the AfterQuery team for their support and review of the data.


Citation

If you use this benchmark in your research, please cite the paper.

Citation (BibTeX)

@misc{bigeard2025fab,
title        = {Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks},
author       = {Bigeard, Antoine and Nashold, Langston and Krishnan, Rayan and Wu, Shirley},
year         = {2025},
month        = may,
howpublished = {Vals AI},
url          = {https://arxiv.org/abs/2508.00828},
}