Partners in Evaluation
Key Takeaways
- Unlike most benchmarks here, the top is not a dead heat: Claude Opus 5 leads at 55.29% all-pass accuracy, a clear ~6-point step ahead of the next model — so on legal research there is a single, meaningfully-best model rather than an interchangeable cluster.
- Under partial-credit scoring, Claude Opus 5 reaches 90.58% weighted pass rate but 55.29% under strict all-pass grading, where every rubric check must pass. The gap shows models often get most of an answer right but fail on one or two required elements.
- Performance varies considerably by practice area. Health and Administrative / Regulatory score highest on average; Family and Immigration are by far the hardest.
- Accuracy doesn’t track tool use: Claude Fable 5 reaches second place with about 15 turns and 40 tool calls per task, while GPT-5.6 Sol uses roughly 47 turns and 173 calls for a similar score, and the weakest performers run far more without the payoff.
- Reconciling conflicting authority is the most reliable failure mode: every model scores lower on questions that require synthesizing across jurisdictions, courts, or regimes than on those that don’t, a drop of 6 to 17 points per model and 11 points when pooled.
Background
Legal Research Bench evaluates AI agents on realistic legal research tasks drawn from diverse areas of US law. Each task requires an agent to research a legal question using a set of tools including case law search, web search, and document retrieval, then produce a well-supported answer.
The benchmark tests whether models can conduct the kind of multi-source research that junior associates and paralegals routinely perform: finding relevant statutes and case law, applying precedent to a fact pattern, and synthesizing across jurisdictions or regulatory regimes. Spanning U.S. federal and state law, the questions require more than simple retrieval. Each demands multiple steps of legal reasoning rather than a single lookup.
The tasks span eight practice areas: Administrative / Regulatory, Business & Commercial, Civil Litigation, Constitutional / Civil Rights, Criminal, Family, Health, and Immigration. This breadth tests general legal reasoning ability across common research workflows.
All questions, gold-standard answers, and rubrics are authored and peer-reviewed by practicing lawyers, drawing on real research questions that arise in legal practice.
Scoring
Each question carries a rubric, written by legal experts. Experts include only items that a correct answer must contain; each rubric item is a required element, not an optional one. Rubrics weight items by importance and range from 1 to 31 items (mean 9.35, mode 10).
Our primary metric is all-pass: a question counts as correct only if the response satisfies every rubric item: 100% if all checks pass, 0% otherwise. We also report a partial-credit weighted score, the share of rubric points earned. All grading is performed by an LLM judge; see Methodology for the harness and judge validation.
Results
The Pareto chart above shows how accuracy trades off against cost and latency across 27 models, with additional metrics such as answer length and sources cited available as alternate axes. Claude Opus 5 leads on strict accuracy at 55.29% at a cost of ~9.79 per task and ~22 minutes, while GPT-5.6 Sol sits third at 48.08% for ~$21.61 per task and ~77 minutes. Among the top three, Opus 5 is both the most accurate and the least expensive, while Fable 5 is the fastest.
Performance by Practice Area
The benchmark spans eight broad areas of U.S. federal and state practice, weighted toward the domains where professional legal research is most often performed. In descending order of coverage: Administrative / Regulatory, Criminal, Business & Commercial, Constitutional / Civil Rights, Health, Civil Litigation, Family, and Immigration. The four largest categories together account for roughly 70% of the benchmark. Because many questions implicate more than one area of law, these categories are not mutually exclusive; the distribution reflects the interconnected nature of legal research, where a single matter frequently draws on multiple bodies of law at once.
Mean all-pass accuracy varies widely across practice areas, from 44.5% in Health and 36.5% in Administrative / Regulatory at the top down to 23.5% in Immigration and 13.6% in Family at the bottom. Business & Commercial, Constitutional / Civil Rights, Civil Litigation, and Criminal cluster between 26.0% and 28.8%.
The chart above shows the best single-model all-pass score in each practice area, with areas ordered by their average score; hover to see the top three models. No one model leads everywhere: Claude Opus 5 tops Administrative / Regulatory (65.2%), Business & Commercial (62.8%), Civil Litigation (63.0%), and Family (40.9%); Claude Fable 5 leads Constitutional / Civil Rights (50.0%); Claude Opus 4.8 leads Criminal (45.6%); GPT-5.6 Sol leads Health (80.0%); and Muse Spark 1.2 leads Immigration (54.5%), with GPT-5.6 Sol and Kimi K3 tied for second (45.5%).
The per-model heatmap below breaks the same data out model-by-model, with practice areas ordered hardest-first (left) to easiest (right). Health and Administrative / Regulatory stay green across nearly every model, while Family is the reddest column, though Claude Opus 5 reaches 40.9% all-pass there. The accuracy spread within a single area is wide: Health ranges from 0% to 80% across models, so practice-area difficulty and model capability compound rather than one dominating the other.
Performance by Question Type
Each task carries two kinds of label: base reasoning types that capture the core legal work a question demands, and overlay flags that mark an added source of difficulty layered on top of that work. A question may involve more than one base type and may carry either, both, or neither flag, so each bar below reports the all-pass rate across all questions that carry that label.
There are three base reasoning types:
- Statutory Interpretation: determining the meaning and scope of statutory text and how courts have construed it.
- Regulatory Framework Interpretation: working out how administrative regulations and agency guidance implement a statute and operate together.
- Doctrinal Rule Reasoning: identifying the controlling legal test and applying precedent to a fact pattern.
Pooled across the evaluated models, Regulatory Framework Interpretation is the most tractable at 37.8% all-pass, while Statutory Interpretation (26.4%) and Doctrinal Rule Reasoning (25.5%) are harder, since both demand reasoning beyond locating the governing text.
The two overlay flags mark complications that can attach to any base type:
- Reconciliation: the question requires synthesizing across multiple authorities, such as comparing jurisdictions, weighing agreement against disagreement among courts, resolving conflicting precedent, or untangling interacting legal regimes.
- Temporal Validity: the answer turns on whether a rule is still good law or on how the rule evolved over time.
Reconciliation is the clearest difficulty signal in the benchmark. At 20.7% all-pass it is the hardest category of all, well below the 28.4% overall rate, and the gap holds for every single model: each one scores lower on reconciliation questions than on the rest. Synthesizing conflicting authority, rather than locating a single rule, is where models break down most reliably. Temporal Validity questions, by contrast, score 27.3%, close to the overall rate, so reasoning about whether a rule still holds is not by itself a major obstacle.
The heatmap below breaks the same categories out model by model, with overlay flags marked in orange.
Weighted vs All-Pass Scoring
Under partial-credit scoring, Claude Opus 5 reaches 90.58% weighted pass, while its strict all-pass accuracy is 55.29%. The gap shows models often satisfy most rubric checks but miss one or two required elements.
We report all-pass as our primary metric because, in legal work, a partially correct answer can be more dangerous than a wrong one: it may read as sound while omitting a critical point. A weighted score above 80% can mask that an answer misses a required element.
Sources Cited
Most models cite more sources on average than the gold-standard answers do, but source quality and relevance, not count, drive the source-check score in grading.
Answer Length
Every model writes far longer answers than the gold-standard references, which average 536 words. Claude Sonnet 4.6 is the most verbose at roughly 2,300 words (over four times the reference length), while GPT 5.4 Mini is the most concise. Length tracks loosely with verbosity rather than accuracy: the longest responses are not the highest-scoring, and grading rewards correct, well-supported reasoning over volume.
Tool Call Analysis
Agents iterate with legal research tools (case-law search, web search, document retrieval, and HTML parsing), storing fetched content in a session database until they submit a final answer. Each task has a three-hour time limit.
web_search is the most-used tool across models, averaging 26 calls per session versus 11 for courtlistener_search. Some models allocate effort more selectively: Claude Opus 4.8 balances case-law and web search in roughly equal proportion, while others run many more total calls, especially web_search, without converting that into higher accuracy. retrieve_information and parse_html_page support extraction from fetched documents once sources are identified.
Model Output Examples
Below are the agents’ full answers to a single representative Business & Commercial question, one of the benchmark’s public sample tasks: a Virginia equipment-leasing dispute that turns on UCC Article 2A finance-lease rules. Switch tabs to compare how each model framed its analysis; models are ordered by overall benchmark accuracy. Claude Sonnet 5 is not shown for this question: its output was blocked by the provider’s content-filtering policy.
Client is a Virginia corporation ("Client") that hauls refrigerated freight out of Rockbridge County. In early 2024, Client needed two new reefer trailers and went to a national manufacturer ("Manufacturer") to pick out a model and negotiate the refrigeration specs. Client then arranged for a Virginia equipment financing company ("Lessor") to buy the trailers ("Trailers") from Manufacturer for $85,000 each and lease them to Client under a written 60-month lease ("Lease") at $2,200 per month per trailer, and Lessor gave Client a complete copy of its purchase agreement with Manufacturer, which included Manufacturer's standard two-year warranty on all mechanical and refrigeration components, before Client signed.
Four months after delivery, both Trailers had compressor failures that left the units unable to hold temperature, and Manufacturer agreed to fix them under the warranty but filed for Chapter 7 two months later, where the trustee confirmed it wouldn't honor outstanding warranty claims. That said, the Lease doesn't include any warranties from Lessor and tells Client to take equipment claims to Manufacturer. It also says that Client's payments are "absolute and unconditional, irrespective of any defense or any right of setoff, counterclaim, or recoupment." Client has stopped paying Lessor. Given these facts, can Client stop its payments to Lessor, and is Lessor responsible for the defective Trailers?
Trajectory Analysis
The chart below traces how each model researched that same question. The top models by accuracy are shown by default; use the selector to add or remove others. Claude Sonnet 5 is omitted here because its output on this question was blocked by the provider’s content-filtering policy. Every tool call is bucketed into a category and laid out in the order it was made, so each bar shows both how long the run was and how the model spent its effort.
The models reach an answer in very different ways. Claude Opus 4.8 works through 38 tool calls, opening with web search to frame the issue, then leaning on document reading and extraction to parse the Virginia statutes and pull specific passages, with a few targeted case-law searches before answer submission. GPT 5.5 takes a more case-law-driven path across 49 calls. Kimi K2.6 runs by far the longest at 174 calls, the vast majority of them web searches, without converting that volume into stronger accuracy. At the other extreme, Gemini 3.1 Pro Preview (02/26) is the most economical at 19 calls, with GLM 5.2 close behind at 37. Every model closes with one submission step.
Methodology
Agents are evaluated on a shared harness with access to five tools. Each task has a three-hour time limit; the agent may iterate with tools until it submits a final answer or the limit is reached.
courtlistener_search: searches the CourtListener database for case law, statutes, and legal documentsweb_search: general web search for secondary sources, regulatory text, and commentaryretrieve_information: queries over previously fetched documents to extract specific passagesparse_html_page: downloads and parses an HTML page from a URLsubmit_final_result: submits the agent’s final answer for evaluation
Grading
All responses are graded by GPT 5.4 as a judge against each task’s rubric (see Scoring for the all-pass metric and rubric structure). We selected this judge because its grades align more closely with a majority vote of 3 human expert reviewers than any single human reviewer does with that majority: GPT 5.4 agrees with the majority vote 87.4% of the time, compared to an average human baseline agreement of 83.4%.
Dataset
The dataset comprises 413 expert-authored, peer-reviewed questions, each paired with a gold-standard answer, authoritative legal sources, and a detailed grading rubric. It is divided into three parts: Public (5 open-source samples), Private Validation (200 samples available for license), and Test (208 samples).
- The Public set and agent harness are fully open and can be accessed here.
- The Private Validation set is available for license. Interested parties are encouraged to contact us directly for access.
- The Test set will remain private. All results reported on this page are based solely on the Test set to prevent overfitting.
The dataset splits were sampled to preserve the distribution of question types, practice areas, and difficulty.
Citation
If you use this benchmark in your research, please cite:
Citation (BibTeX)
@misc{valsai2026legalresearch,
title = {Legal Research Bench: Evaluating Agents on US Legal Research Tasks},
author = {Vals AI},
year = {2026},
month = jun,
howpublished = {Vals AI},
url = {https://github.com/vals-ai/legal-research-bench},
}