Proprietary

Finance Agent

Updated 8/14/2026

Read research paper

Evaluating agents on core financial analyst tasks

Finance Agent v2Core financial analyst tasks
ACCURACY

Key Takeaways

  • Muse Spark 1.2 leads Finance Agent v2 with 60.60% accuracy. Claude Opus 5 follows at 58.63%, with Gemini 3.5 Flash at 57.86%.
  • Models are able to handle simple retrieval tasks, but still struggle to perform reliably on harder, multi-step financial work that relies on precise numbers and specific industry convention.
  • Muse Spark 1.2 is the only model above 60% with Partial Credit. Under stricter All-Pass scoring, it reaches 50.88%, while every other model remains below 48%.
  • There is substantial room for improvement: no single model leads all nine question categories, and the hardest categories remain Financial Modeling and Precedents, where the category leaders reach only 34.52% and 36.37%, respectively.
  • One of Opus 5’s 1,350 run-task results used Claude Opus 4.8 as a refusal fallback. That task already scored 0.0, so the published accuracy is unchanged.

Background

Finance Agent v2 builds on Finance Agent v1.1 with 927 expert-reviewed questions across public, private validation, and held-out test splits. Dataset design, splits, and access are covered in Methodology.

The benchmark tests a model’s ability to perform the work of entry-level financial analysts — answering difficult questions on public company filings.

Automating analyst-grade financial research is valuable because the work is expensive, repetitive, and time-sensitive: analysts often need to move from filings and transcripts to a defensible model or investment memo under deadline. The work is also difficult because the answer depends on finding the right source, applying the right finance convention, and carrying precise intermediate numbers through several steps.

Finance Agent v2 measures how well AI can aid analysts in automating busy work so they can focus on higher-leverage tasks.


Results

The Pareto chart above illustrates how model accuracy trades off against cost and latency. Muse Spark 1.2 leads at 60.60% while also being the cheapest and fastest of the top three, at about a seventh the cost of Claude Opus 5, which follows at 58.63%. Gemini 3.5 Flash is close behind at 57.86% while costing about half as much as Opus 5 and running nearly twice as fast.

The mid-tier models sit only a few points below the leaders at a fraction of the cost, suggesting the accuracy premium at the frontier is modest relative to the price gap.

Where Models Stand

Category LeadersHover for top three
General Quantitative
Gemini 3.7 Flash81.8%
Earnings Analysis
Gemini 3.7 Flash79.1%
General Qualitative
Grok 4.677.8%
Market Analysis
Gemini 3.7 Flash74.5%
Disclosure Analysis
Claude Opus 571.3%
Adjustments
Muse Spark 1.256.3%
Comparables
Gemini 3.7 Flash50.3%
Precedents
Gemini 3.5 Flash36.4%
Financial Modeling
Muse Spark 1.234.5%
Categories ordered by leader accuracyBars scaled 0–100% accuracy · gridlines every 25%

The retrieval and summarization categories cluster at the top: General Quantitative, Earnings Analysis, General Qualitative, Market Analysis, and Disclosure Analysis all have category leaders above 70%. Adjustments tops out at 56.29%, while Comparables reaches 50.24% as cross-document reconciliation gets harder. Financial Modeling and Precedents remain the hardest categories, topping out at 34.52% and 36.37%, respectively.


Partial Credit vs All-Pass
PARTIAL CREDITALL-PASS

When measuring models on their ability to perfectly answer questions, scores fall substantially across the leaderboard compared with allowing partially complete responses.

The strongest model reaches 50.88% All-Pass accuracy, the next group sits in the mid-to-high 40s, and mid-tier models land in the low 30s. Models are still unable to execute financial tasks perfectly, often missing core details that real analysts would include.


Tool Calls Analysis
TURNSTOOL CALLSERRORS

The tool call analysis reveals how many steps the models take to produce an answer. Interestingly, the top performers also register relatively high numbers of tool calls. This suggests that their comparatively strong performance is partly due to leveraging the tools provided.

Tool Usage
4/49 models

The radar graph shows the per-model breakdown of tool calls across the six tools.

Calculator is the dominant tool for most frontier models. Higher-performing models tend to spend more of their tool budget on numerical resolution after finding the relevant source material. Lower-performing models lean harder on web_search and retrieve_information, making more exploratory calls that do not translate into correct answers. Across the models shown, price_history accounts for at most 4.22% of tool calls, matching its narrower role in Market Analysis questions.


Model Output Examples

Question

Did NYSE: CNC have to pay a rebate to policyholders in fiscal year 2025 or fiscal year 2020? If not, was the company closer to owing customers a rebate in 2025 or 2020? What were the medical loss ratios in each year?

### Summary of Findings

#### 1. Did NYSE: CNC have to pay a rebate to policyholders in fiscal year 2025 or fiscal year 2020?
**Yes, in both fiscal year 2020 and fiscal year 2025**, Centene Corporation (NYSE: CNC) had obligations to pay/accrue rebates and returns of premium to policyholders, customers, and government programs (such as state Medicaid agencies and CMS) under minimum Medical Loss Ratio (MLR) / Health Benefits Ratio (HBR) requirements and contractual risk-sharing/profit-sharing provisions:

* **Fiscal Year 2020:**
* **ACA Minimum Medical Loss Ratio Payable:** Centene reported an ACA Minimum MLR liability of **$238 million** as of December 31, 2020 (compared to **$367 million** as of December 31, 2019).
* **Return of Premium Payable:** Centene recorded a total current liability for return of premium payable on its Consolidated Balance Sheets of **$1,458 million** as of December 31, 2020 (compared to **$824 million** as of December 31, 2019).
* **Reduction to Premium Revenues:** Centene recorded **$86 million** as a direct reduction to premium revenues in 2020 as a result of minimum HBR and other return of premium programs related to prior-year claim developments.
* **ACA Risk Corridor Settlement:** In Q3 2020, Centene recorded its ACA risk corridor settlement net of minimum medical loss ratio paybacks.

* **Fiscal Year 2025:**
* **ACA Minimum Medical Loss Ratio Payable:** Centene reported an ACA Minimum MLR liability of **$294 million** as of December 31, 2025 (compared to **$688 million** as of December 31, 2024).
* **Return of Premium Payable:** Centene recorded a total current liability for return of premium payable on its Consolidated Balance Sheets of **$1,592 million** as of December 31, 2025 (compared to **$2,008 million** as of December 31, 2024).
* **Reduction to Premium Revenues:** Centene recorded **$93 million** as a reduction to premium revenues in 2025 as a result of minimum MLR and other return of premium programs related to prior-year claim developments.

---

#### 2. Was the company closer to owing customers a rebate in 2025 or 2020?
**Centene was significantly closer to owing customers a rebate in 2020 than in 2025.**

* **Mechanism of MLR Rebates:**
Under the Affordable Care Act (ACA) and state/federal Medicaid managed care regulations, health insurers are required to spend a minimum percentage of premium revenues on medical care and healthcare quality improvement activities (typically **80%** for individual and small group commercial plans, and **85%** for large group commercial and Medicaid plans).
* If the Medical Loss Ratio / Health Benefits Ratio falls **below** these statutory/contractual thresholds, the insurer must issue a **rebate / return of premium** for the shortfall to policyholders or state agencies.
* Therefore, a **lower** MLR/HBR brings an insurer closer to (or further below) the minimum MLR threshold, increasing rebate obligations.
* **Comparison:**
* In **2020**, Centene's consolidated MLR/HBR was **86.2%** (86.22%), which was only **1.22 percentage points** above the 85% threshold and **6.22 percentage points** above the 80% threshold. The lower ratio was driven by deferred and reduced medical utilization during the COVID-19 pandemic.
* In **2025**, Centene's consolidated MLR/HBR rose to **91.9%** (91.92%), which was **6.92 percentage points** above the 85% threshold and **11.92 percentage points** above the 80% threshold, driven by elevated medical costs and utilization in Medicaid, Marketplace risk adjustment dynamics, and Inflation Reduction Act (IRA) changes in Medicare PDP.
* Consequently, Centene's medical spending as a percentage of premiums was much lower and much closer to the statutory minimum rebate floors in **2020**.

---

#### 3. What were the Medical Loss Ratios / Health Benefits Ratios in each year?

Centene reports its Medical Loss Ratio as the **Health Benefits Ratio (HBR)**, defined as medical costs as a percentage of premium revenues (excluding separately billed premium taxes):

* **Fiscal Year 2020:**
* **Reported HBR:** **86.2%**
* **Calculated HBR:** $\frac{\$86,264\text{ million}}{\$100,055\text{ million}} = \mathbf{86.22\%}$
* *Underlying Financials:* Medical Costs = $86,264 million; Premium Revenues = $100,055 million.

* **Fiscal Year 2025:**
* **Reported HBR:** **91.9%**
* **Calculated HBR:** $\frac{\$157,702\text{ million}}{\$171,556\text{ million}} = \mathbf{91.92\%}$
* *Underlying Financials:* Medical Costs = $157,702 million; Premium Revenues = $171,556 million.

---

### Sources

```json
{
"sources": [
{
"url": "https://www.sec.gov/Archives/edgar/data/1071739/000107173926000049/cnc-20251231.htm",
"name": "Centene Corporation Form 10-K for the Fiscal Year Ended December 31, 2025"
},
{
"url": "https://www.sec.gov/Archives/edgar/data/1071739/000107173921000039/cnc-20201231.htm",
"name": "Centene Corporation Form 10-K for the Fiscal Year Ended December 31, 2020"
}
]
}
```

60

Missed checks:

• Response states that the company did not owe a rebate in 2025

• Response states that the company did not owe a rebate in 2020

The question above is a General Quantitative Analysis task: deciding whether Centene owed a rebate to policyholders in either fiscal year based on the medical loss ratio threshold, and reporting the underlying MLRs.


Trajectory Comparison

The visualizations below show how a frontier model and a smaller, cheaper model approached the same Financial Modeling question of building a DCF model. Each row is one agent turn; bar width is wall-clock time per turn (model reasoning + tool execution).

Claude Opus 4.7 trajectory (passed most checks):

Claude Opus 4.7 trajectory

The frontier model worked through the question in 12 turns: pull market data, fetch and parse the two filings, run retrieve_information to extract the inputs needed for the DCF, then step through bursts of parallel calculator calls to compute projections, terminal value, and discounted cash flows before submitting.

GPT 5.4 Nano trajectory (zeroed on most checks):

GPT-5.4-nano trajectory

The smaller model needed 34 turns to reach an answer of similar shape — nearly 3x as many. After the retrieval phase, it hammers the calculator one operation at a time for the next 27 turns rather than fanning out parallel calls, and the resulting numbers land close to the rubric’s targets but not exactly enough to be rewarded.


Methodology

Agents are all evaluated on a shared default harness with access to six tools: edgar_search (SEC EDGAR API), web_search, parse_html_page (download an HTML page), retrieve_information (query over fetched HTML), calculator, and price_history.

Finance Agent v2 harness: the LLM iterates with search, process, and fetch tools, storing fetched documents in a database it can query on demand

Time Limit

Each task has a two-hour time limit. An answer must be produced within the time limit, otherwise the task is scored as a zero. The time limit was chosen empirically by measuring how long current frontier models need to converge on these tasks and setting the limit comfortably above that ceiling. A fixed time budget also allows us to normalize and compare performance across different agent implementations.

Grading

Each question is composed of weighted checks, and a subset are flagged as dealbreakers, load-bearing facts or numbers, that are required for a satisfactory answer. Failing any dealbreaker means the answer receives no credit for that question, regardless of the remaining content in the response. Two metrics are reported:

  • Partial Credit (primary): the dealbreaker-gated, severity-weighted average of per-check scores. A response with a correct dealbreaker but a few peripheral misses scores below 100%; a response with any failed dealbreaker scores 0%.
  • All-Pass (secondary): 100% only if every check passes, 0% otherwise.

All responses are graded by a three-judge LLM jury consisting of three frontier models: GPT-5.4, Gemini-3.1-Pro, and Claude Sonnet 4.6.

Changes from v1.1

  • Step up in difficulty. Even the best model reaches only 60.60% with Partial Credit and 50.88% under All-Pass grading. Questions require connecting data across multiple documents, tighter numeric precision, and answers are expected to include insights that skilled analysts would include but are not explicitly stated in the question.
  • Expanded taxonomy. The taxonomy was reorganized around real analyst workflows rather than retrieval tiers. v1.1’s easier retrieval-focused buckets (Quantitative Retrieval, Qualitative Retrieval, Numerical Reasoning, Complex Retrieval, Beat or Miss, Trends) are replaced with Comparables, Precedents, Earnings Analysis, Disclosure Analysis, and a split between General Qualitative and General Quantitative analysis.
  • New grading mechanism. Dealbreaker-gated Partial Credit replaces v1.1’s flat per-question score, and All-Pass is reported alongside it as a strict secondary metric.
  • Stricter numeric tolerances. Tighter thresholds on rounding and precision drift. Answers that previously passed under v1.1’s looser tolerance now fail.
  • Expanded harness. Adds calculator and price_history on top of the v1.1 tools.
  • Multi-run aggregation. Every model is run three times; reported scores are mean-of-runs with standard error of the mean.
  • Expanded test set. Larger held-out test split for tighter measurement.

Question Design

Questions target the analytical depth expected of a 2nd or 3rd-year investment banking analyst. Each question was designed to satisfy four criteria:

  • Determinism. A single, unambiguous correct answer with no room for competing interpretations.
  • Multi-source synthesis. Answers require chaining information across multiple filings or data sources rather than a single lookup.
  • Domain specificity. Questions require implicit industry knowledge that follows sector convention rather than explicit instruction.
  • Forensic precision. Critical information is frequently buried in footnotes, MD&A caveats, or accounting policy disclosures.

Dataset

The dataset is divided into three parts: Public (27 open-source samples), Private Validation (450 samples available for license), and Test (450 samples).

  • The Public set and agent harness are fully open and can be accessed here.
  • The Private Validation set is available for license. Interested parties are encouraged to contact us directly for access.
  • The Test set will remain private. All results reported on this page are based solely on the Test set to prevent overfitting.

The dataset splits were sampled to preserve the distribution of question categories and difficulty.

Question Taxonomy

Finance Agent v2 organizes questions into nine analytical categories reflecting real equity-research workflows.

General Qualitative Analysis

Summarization and comparison of fundamental filing sections: business model, risk factors, MD&A, and standard disclosures across companies.

Compare Walmart, Costco, and Target’s capital allocation priorities across capex, dividends, share repurchases, and debt management.

General Quantitative Analysis

Extraction and calculation of reported financials such as revenue growth, CAGR, leverage ratios, and executive compensation — often requiring verification against restated historicals.

Compare Home Depot and Lowe’s FY2024 inventory efficiency and calculate the difference in days inventory outstanding.

Market Analysis

Relative trading performance, total shareholder return, and how news cycles or guidance shifts drive stock volatility relative to sector indices.

Measure Sun Communities’ stock reaction after the announced sale of Safe Harbor Marinas, then relate the move to the company’s stated use of proceeds.

Comparables

Building trading comps tables, calculating EV multiples, and normalizing enterprise value across peers by adjusting for off-balance-sheet items buried in footnotes.

Rank major U.S. banks by excess CET1 ratio relative to their regulatory minimums.

Precedents

Analyzing M&A transaction multiples from S-4 filings and target financials, with industry-specific EBITDA normalization (e.g. exploration expense add-backs in Oil & Gas).

Extract enterprise values and EV/EBITDA multiples for recent industrial distribution acquisitions and rank the transactions by pre-synergy multiple.

Adjustments

Bridging GAAP to non-GAAP or pro forma figures by reconciling SBC, acquired intangible amortization, and other non-cash items across the P&L and cash flow statement.

Reconcile Honeywell’s GAAP operating income to segment profit across annual releases and identify newly introduced adjustment categories.

Earnings Analysis

Comparing reported results against consensus estimates and prior guidance, including non-GAAP reconciliations across consecutive quarterly reporting cycles.

Compare Rapid7’s Q3 2025 actuals against prior revenue, non-GAAP operating income, and ARR guidance.

Disclosure Analysis

Tracking shifts in MD&A language, KPI definitions, and segment reporting methodology across multiple annual filings, then restating prior periods to reflect the new format.

Track Boeing’s segment reporting and 787 cost-recovery disclosures across FY2022-FY2024 10-K filings.

Financial Modeling

Multi-step frameworks including DCF/NPV, LBO, and M&A accretion/dilution models built from historical ratios extracted from primary filings.

Assess whether Ralph Lauren could justify a distressed acquisition of Capri under stated synergy, margin, and valuation assumptions.


Acknowledgements

We would like to thank Andrew Schettino and all of the financial experts who worked on Finance Agent v2.


Citation

If you use this benchmark in your research, please cite the paper.

Citation (BibTeX)

@misc{bigeard2025fab,
title        = {Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks},
author       = {Bigeard, Antoine and Nashold, Langston and Krishnan, Rayan and Wu, Shirley},
year         = {2025},
month        = may,
howpublished = {Vals AI},
url          = {https://arxiv.org/abs/2508.00828},
}