Proprietary

Tax Agent Bench

Updated 9/4/2026

Evaluating agents on research-grade US corporate tax questions

Tax Agent BenchAgentic US corporate tax research
ACCURACY

Key Takeaways

  • The top is a close race: Claude Fable 5.1 leads at 77.64%, with Claude Opus 5 (75.06%) and GLM 5.3 (73.09%) within five points. No model clears 80% overall, so even the best agents still miss required elements on a meaningful share of corporate tax questions.
  • Forms & Filings and Controversy & Precedence Analysis are the hardest categories, with no model above 74%: producing filing-ready work product and determining which rules and rulings apply to a given task remain harder than computing the right figure. Numbers & Calculations is the easiest category for nearly every model (top score 87.97%).
  • Top performing agents shift by category. GLM 5.3 leads Rule & Source Lookup (81.29%) and Controversy & Precedence Analysis (71.08%) despite ranking third overall, and GPT-5.6 Sol is second on Forms & Filings despite ranking seventh.

Background

Tax Agent Bench evaluates AI agents on realistic, research-grade US tax questions, with a focus on corporate tax. Rather than static Q&A, tasks reflect the multi-step research process a tax practitioner actually follows: identifying the governing authority, synthesizing statutes, regulations, IRS guidance, and forms, applying the rules to a fact pattern, and producing an accurate, well-supported answer.

Each task is a single written question. The agent researches it with a set of tools over as many steps as it needs, then submits one final written answer with its citations. Only that final answer is graded, against the question’s expert rubric and a check that its cited sources are real; the intermediate trajectory (the sequence of tool calls) is recorded and shown on this page for analysis but is not scored.

The questions center on the work of an enterprise tax department: federal corporate income tax (roughly 60% of the test set), with the remainder spanning partnership tax, state and local tax (SALT), transfer pricing, and corporate employment tax. Fact patterns involve C corporations, consolidated groups, reorganizations, reportable-transaction disclosure, penalties and procedure, and multi-year or multi-jurisdiction comparisons.

The benchmark tests whether models can conduct the kind of research that tax associates perform daily, where a single matter frequently requires chaining several lookups and reasoning steps, and where an early error in a threshold, date, or citation propagates to the final answer.

All questions, gold-standard answers, and rubrics are authored and reviewed by tax professionals.


Scoring

Each question carries a rubric of required checks written by tax experts (3 to 89 checks per question, mean 23.6, mode 10). Every check carries a weight from 1 to 3 reflecting how much it matters to the answer, and roughly 40% of checks are marked must-pass: elements a correct answer cannot omit, such as identifying the controlling Code section or the required form.

Our primary metric, Accuracy, is a weighted partial-credit score. For each question, the rubric score is the share of weighted rubric points earned across its checks, set to 0 if any must-pass check fails. Citation quality is the fraction of sources cited in the answer that resolve to a real, authoritative document. The task score combines the two as

task score = rubric score × (0.7 + 0.3 × citation quality)

so an answer whose citations all resolve keeps its full rubric score, one whose citations all fail keeps 70% of it, and one with half its citations resolving keeps 85%. Citation quality can therefore reduce a task score by at most 30%, and it cannot rescue an answer that failed a must-pass check. Scores are averaged over all 193 questions of the private test set with every task weighted equally; a question the agent did not answer within the time limit scores 0. We also report All-Pass, a strict secondary metric: the proportion of questions on which the model satisfied every rubric check. All grading is performed by an LLM judge; see Methodology for the harness and judge.


Results

Unless otherwise noted, all results on this page are computed on the private 193-question test set. The Pareto chart at the top of the page shows how accuracy trades off against cost and latency across 18 models. Claude Fable 5.1 leads at 77.64% for ~$13.18 per task and ~29 minutes. Claude Opus 5 follows at 75.06% for ~$5.14 per task and ~17 minutes, the fastest of the top three and less than half the cost of the leader, while GLM 5.3 is third at 73.09% for ~$1.80 per task and ~51 minutes. Muse Spark 1.3 is fourth at 71.93% for ~$0.32 per task and ~8 minutes, cheaper and faster than every model above it, making it the strongest budget option ahead of Grok 4.6 (70.79%, ~$0.98, ~11 minutes); Gemini 3.8 Flash (66.77%, ~$0.86, ~3 minutes) sits eighth overall and fourth on Fact-Pattern Analysis (67.95%), while its predecessor Gemini 3.7 Flash answers in about 1.5 minutes at 57.67%.

Performance by Category

Questions fall into one of six categories. The first three are bounded research tasks with a single locatable answer: the right rule, the right figure, or the right form. The latter three require reasoning across procedure, time, and interacting facts, where the answer depends on weighing authority rather than retrieving it. The counts below are for the 193-question test set; the validation set follows the same distribution (31/30/30/30/30/42):

  • Rule & Source Lookup (30 test questions): identifying the governing tax rule, threshold, or requirement and locating the most authoritative source among statutes, regulations, forms, and IRS guidance.
  • Numbers & Calculations (30): determining applicable percentages, dollar thresholds, and due dates, and performing multi-step computations where an early error propagates to the final answer.
  • Forms & Filings (30): identifying filing requirements, disclosure obligations, and information returns, and producing work product such as memos or completed forms.
  • Controversy & Precedence Analysis (30): audits, penalties, appeals, statutes of limitation, elections, and procedural requirements, including determining which rules and rulings apply to a given task.
  • Current & Temporal Analysis (30): recent law changes, new forms, and recent guidance, or comparing rules across tax years, jurisdictions, or alternative fact patterns.
  • Fact-Pattern Analysis (43): applying one or more tax rules to a specific fact pattern, from a single rule applied to a clearly stated set of facts to scenarios where multiple concepts interact and require synthesis across sources.

Use the task selector on the results table to view each category. Forms & Filings and Controversy & Precedence Analysis are the hardest categories, with best scores of 73.18% (Claude Fable 5.1) and 71.08% (GLM 5.3) respectively. Numbers & Calculations is the easiest: the top five all exceed 83%, and even the weakest model reaches 54.9%.

No model leads everywhere. Claude Fable 5.1 tops Fact-Pattern Analysis (79.29%), Current & Temporal Analysis (81.21%), Numbers & Calculations (87.97%), and Forms & Filings (73.18%); GLM 5.3 leads Rule & Source Lookup (81.29%) and Controversy & Precedence Analysis (71.08%). The spread within a category is wide: on Rule & Source Lookup, GPT 5.5 scores 44.23% while GLM 5.3 scores 81.29%, so locating the right authority is far from solved even for frontier models.

All-Pass: Fully Correct Answers

Accuracy awards partial credit for each rubric check an answer passes, weighted by importance. All-Pass asks the stricter question a client would: was the answer fully right? A question counts only if every rubric check passes.

Accuracy vs All-Pass
ACCURACYALL-PASS

Under All-Pass, no model reaches 50%. Claude Fable 5.1 leads at 49.22%, followed by Claude Opus 5 (45.60%), GLM 5.3 (39.90%), and Muse Spark 1.3 (36.79%); every other model is below 34%. The gap between the two metrics is between 26 and 43 points for every model, so the typical answer earns most of its rubric points but still omits at least one required element. The ranking also shifts: Kimi K3 (33.68%) edges past Grok 4.6 (33.16%) on All-Pass despite trailing it by two points on Accuracy, and GPT 5.5 drops to 21.76%, below Gemini 3.7 Flash, because its answers more often miss one check outright. Switch the metric selector above to All-Pass to view it overall or per category.

Tool Use

The harness exposes five tools: web_search, lookup_authority, fetch_document, retrieve_information, and calculator (each is defined under Methodology). Every model uses lookup_authority, which returns statutory and regulatory text by citation, on every question, but the overall research style differs sharply: the three leaders make it their most-used tool, while search-heavy agents such as Grok 4.6, Kimi K3, and the OpenAI 5.6 checkpoints issue more web_search calls than authority lookups.

Tool Calls Analysis
TURNSTOOL CALLSERRORS

The three OpenAI 5.6 checkpoints make far more calls than anyone else, 113 tool calls per question for GPT-5.6 Sol, about 100 for GPT-5.6 Terra, and 91 for GPT-5.6 Luna, roughly double the 41 to 49 of the three leaders. The extra calls are mostly web searches (43, 37, and 33 per question versus 6 to 8 for the two Claude leaders) and do not translate into accuracy: they score 61% to 68%. Grok 4.6 and Kimi K3 are also search-heavy, with 21 and 25 web searches per question, while the top three models instead spend their calls on lookup_authority (10 to 16 per question) and on fetching and querying primary documents. Errors during tool calls are rare across the board, between 0.2 and 3.5 per question, and are highest for the OpenAI 5.6 checkpoints, so the extra volume also carries more failed calls. Across all 18 models, 3.8% of tool calls fail (6,192 of 164,123), and the failures follow a consistent pattern: about three-quarters are fetch_document calls that hit a page the harness could not read, mostly URLs that return 404 or sites that block automated access, with a smaller share of PDFs that yield no text or time out; 13% are lookup_authority requests for a citation that does not resolve, most often a regulation section number that does not exist in that form (for example a call for Treas. Reg. §1.338 without the subsection); and 9% are retrieve_information queries that name no saved document or one the agent never stored, nearly half of them from Inkling.

Tool Usage
18/18 models

The radar chart shows the per-tool mix. The Claude models and GLM 5.3 have the most balanced profiles, pairing authority lookups with fetch_document and retrieve_information to read full regulations and rulings, which is what the Controversy and Fact-Pattern categories require. Gemini 3.7 Flash is the most economical, averaging 20 tool calls and only 1.7 document fetches per question, which explains its 1.5-minute latency but also its 57.7% accuracy. Muse Spark 1.3 shows that a light footprint need not cost accuracy: it averages 24 calls per question, about the same as Inkling and MiniMax-M3 and fewer than every other model except Gemini 3.7 Flash, yet ranks fourth overall. calculator is used between roughly two and eight times per question by every model, most heavily by the OpenAI 5.6 checkpoints.

Answer Length

Tax practitioners expect a research memo, not a one-line answer, and the rubrics reward supporting analysis and citations. The chart below plots each model’s mean final-answer length in words against its Accuracy; the reference line marks the mean length of the expert-written reference answers in the dataset, 601 words.

Answer Length by Model
ANSWER LENGTHGOLD STANDARD (601 WORDS)

The three most accurate models also write the longest answers: Claude Opus 5 averages 2,757 words, GLM 5.3 2,434, and Claude Fable 5.1 2,366, versus 950 to 2,000 for the rest of the field. Length is not sufficient on its own, however: Inkling writes 1,809 words for a 43.5% score, about the same length as Grok 4.6 (1,769 words, 70.8%). The shortest answers come from GPT-6 Astra (956 words) and GPT 5.5 (1,101 words), whose All-Pass rates (20.7% and 21.8%) are the two lowest, consistent with answers that omit required elements rather than getting them wrong. Every model writes far more than the experts did: the shortest model average (956 words) is about one and a half times the 601-word reference mean, and the longest (2,757) is more than four times it. The reference answers satisfy every rubric check in a fraction of the space, so the extra length is not what the rubrics require, and no model is close to the expert length: the verbosity is a property of the agents, not of the task.

Trajectories on a Public Question

To show how the agents actually work through a problem, the chart below traces the tool-call sequence of all 18 models on public question P-004, a Current & Temporal Analysis task in State & Local Tax: for tax year 2025, how do New York’s and California’s pass-through entity taxes (PTET) differ in how the credit is allocated to an individual partner, and is that credit refunded if it exceeds the partner’s state income tax for the year? Each bar is one tool call, in order, colored by tool. For the 14 models run on the full dataset the trajectories come from the same runs that produce the scores above, where this question appears under its full-dataset ID; Gemini 3.8 Flash, GPT-6 Astra, GPT-5.6 Sol, and Muse Spark 1.3 were run on the public split separately from their test-set runs.

Trajectory by Model
Claude Fable 5.1 (52 steps)
Claude Opus 5 (55 steps)
GLM 5.3 (60 steps)

Only Grok 4.6 passes all 18 checks, and it does so with one of the longest trajectories (88 calls, 59 of them web searches). At the other extreme, Muse Spark 1.3 scores 79.4% in 19 calls, the shortest trajectory of any model here, clearing every must-pass check and dropping four lower-weight details. The overall leader Claude Fable 5.1 scores 91%, clearing every must-pass check and missing two lower-weight points (the New York statutory citation and one category of California owner that cannot claim the credit). Kimi K3, GPT-5.6 Luna, and GPT 5.5 each land at 76.5%, getting the refundability answer right for both states but dropping the same cluster of supporting details: the governing statutes, the California claim form, what the electing entity must report, and which owners California excludes. The remaining twelve, including Claude Opus 5 and GLM 5.3, score 0. Eleven of them fail the same must-pass check, how New York treats a partner who receives PTET credits from more than one electing entity, and Inkling fails a different one, which California tax the credit offsets. Effort does not decide it: GPT-5.6 Terra makes 95 calls and scores 0, while GPT 5.5 reaches 76.5% in 40 and GPT-6 Astra scores 0 in the same 40, and Muse Spark 1.3 beats both in half as many calls. Research style differs as much as length: Grok 4.6, Gemini 3.8 Flash, Gemini 3.7 Flash, and GPT-5.6 Terra lean on web_search for more than half their calls, while Claude Fable 5.1, Claude Opus 5, Kimi K3, and DeepSeek V4 Pro 0813 work almost entirely through fetch_document and retrieve_information. On a single question with a hard threshold, the overall leaderboard is not predictive; the aggregate ranking emerges only across the full 193 questions.

Question

For calendar tax year 2025, how do the tax credit allocation and liquidity mechanics differ between New York State's pass-through entity tax (PTET) and California's pass-through entity tax (PTET) regarding whether the resulting credit passed through to an individual partner is fully refundable if it exceeds that partner's personal state income tax liability for the year?

## Short answer

For calendar tax year 2025 the two regimes reach opposite results on the refundability question:

- **New York State PTET (Tax Law Art. 24‑A; credit under Tax Law § 606(kkk))** – the partner's credit is **fully refundable**. Tax Law § 606(kkk)(4): "If the amount of the credit allowable pursuant to this subsection for any taxable year exceeds the tax due for such year pursuant to this article, the excess shall be treated as an overpayment, to be credited or refunded, without interest." The Department restates this in TSB‑M‑21(1)C,(1)I and its PTET FAQ ("the excess credit will be refunded without interest"; the PTET credit "is fully refundable").
- **California PTE elective tax (R&TC §§ 19900–19906; credit under R&TC § 17052.10)** – the partner's credit is **nonrefundable**. It is allowed only "against the 'net tax,' as defined in Section 17039" (§ 17052.10(a)); any excess "may be carried over to reduce the 'net tax' in the following taxable year, and succeeding four years, if necessary, until the credit is exhausted" (§ 17052.10(c)). The 2025 FTB 3804‑CR instructions state flatly: "This credit is not refundable and cannot be assigned," and unused credit "may be carried over to reduce the tax for five years or until exhausted, whichever occurs first … In no event can the credit be carried back." Credit still unused after the fifth carryover year simply expires; the FTB never pays it out in cash or with interest.

Everything else about the two credits' allocation and cash‑flow mechanics flows from that difference.

## 1. New York State – allocation and liquidity mechanics (TY 2025)

**Who gets the credit and how much.** The electing partnership pays PTET on its "PTE taxable income" at graduated rates (6.85% up to $2 million; 9.65%, 10.30% and 10.90% brackets above that) (TSB‑M‑21(1)C,(1)I). Each *direct* partner subject to Article 22 (individuals, trusts, estates) is entitled to a credit "equal to the partner's, member's or shareholder's direct share of the pass‑through entity tax" (§ 606(kkk)(2)); credits from several electing entities are summed (§ 606(kkk)(3)). No partner consent is required; corporate and partnership partners get nothing and cannot pass a credit down (TSB‑M; FAQ). For partnerships the entity computes separate resident and nonresident PTET credit pools and allocates within each pool by profit‑and‑loss percentage (TSB‑M). Two statutory caps: no credit unless the entity identified the partner on its Art. 24‑A annual return under § 865, and the credit "shall not exceed the direct share of pass‑through entity tax reported by such electing partnership … on the entity's return" (§ 606(kkk)(5)). Total credits reported may not exceed the PTET actually paid (TSB‑M).

**Where it is claimed (2025 forms).** The partner files Form IT‑653 (which cites Tax Law §§ 606(kkk) and 1310(g)) with Form IT‑201/IT‑203/IT‑205; the IT‑653 line 3 total is entered with code 653 on 2025 Form IT‑201‑ATT, Part 1, Section D, line 12 ("Other refundable credits"), flows through lines 13/14 to line 18, and then to Form IT‑201, line 71 (IT‑203‑ATT line 12 for nonresidents; IT‑205 line 33 for fiduciaries) (Form IT‑653‑I; 2025 Form IT‑201‑ATT). Because it sits in the refundable‑credit/payments section, it "offsets all taxes computed and reported on" Forms IT‑201, IT‑203 and IT‑205, is applied after nonrefundable credits such as the § 620 resident credit, and any excess "will be refunded without interest" (PTET FAQ) or, at the partner's election, applied to next year's estimated tax (Form IT‑653‑I). The credit may not be claimed on group nonresident returns IT‑203‑GR/IT‑203‑S (TSB‑M).

**Add‑back.** The partner must add back an amount equal to the § 606(kkk) credit claimed (Tax Law § 612(b)(43)(A); Form IT‑225) – so the federal deduction taken by the entity does not also reduce New York taxable income.

**Entity‑level cash flow for 2025.** Election online by March 15 (March 17, 2025, because the 15th was a Saturday); irrevocable after the first estimated‑payment due date. Mandatory quarterly PTET estimates of at least 25% of the required annual payment (lesser of 90% of current‑year PTET or 100% of prior‑year PTET) due March 15, June 15, September 15 and December 15, 2025 (next business day if a weekend: March 17 and June 16, 2025), with Art. 22 penalties/interest for underpayment; annual PTET return and any balance due March 15, 2026 (March 16, 2026), six‑month filing extension available but no extension to pay (PTET webpage; TSB‑M). Entity overpayments cannot be carried forward or transferred to partners' accounts; they are refunded by check to the entity ("No transfer of tax payments to or from PTET is allowed … it cannot transfer payments between related entities or different tax types, including individuals") (PTET webpage; FAQ).

**Partner‑level cash flow.** The partner cannot treat PTET estimates as his or her own IT‑2105 payments, but "Taxpayers may take the PTET credit into account when computing their estimated income tax requirements for the year" (PTET FAQ). The credit must be claimed "for the same tax year the PTET annual return is filed, regardless of when the PTET is paid," and a partner who filed early must amend (FAQ). Net result: any mismatch between the entity‑level PTET (graduated rates on entity income) and the partner's own Article 22 liability is returned to the partner in cash (no interest) with the 2025 return filed in 2026 – the partner bears no "trapped credit" risk.

## 2. California – allocation and liquidity mechanics (TY 2025)

**Governing version.** For taxable years beginning on or after January 1, 2021 and before January 1, 2026, the elective tax is imposed by R&TC § 19900(a)(1) and the credit by § 17052.10 (as amended by SB 132, Stats. 2025, ch. 17, § 7). The successor provisions added by SB 132 for 2026–2030 (§ 17052.11 and Part 10.4.1, §§ 19910 et seq.) do not apply to calendar 2025 (§ 17052.11(a)); I note their differences below.

**Who gets the credit and how much.** The entity pays 9.3% of "qualified net income," i.e., the pro rata/distributive shares and guaranteed payments of the *consenting* qualified taxpayers (§ 19900(a)(1)–(2), (c)(1)). Only a consenting individual, fiduciary, estate or trust (or its disregarded SMLLC) is a "qualified taxpayer"; partnerships, corporations and nonconsenting owners receive no credit (§ 17052.10(b)(3); FTB 3804‑CR instructions). The credit is a flat "qualified amount" = 9.3% of the qualified taxpayer's share of QNI (§ 17052.10(b)(2)) – it is not a "direct share of tax paid" concept and is "not a pass‑through item," though it is shown on the owner's Schedule K‑1 (FTB 3804‑CR instructions). A disallowance because the June 15 payment was not timely made, because payments exceed the computed tax, or because no valid election was made is treated as a mathematical error assessable under § 19051 (§ 17052.10(e)).

**Where it is claimed and why excess cannot be refunded (2025 forms).** The partner attaches FTB 3804‑CR to Form 540/540NR/541 and claims the credit with credit code 242 on Form 540 lines 43–45 (Schedule P (540), Part III if more than two credits or if Schedule P is otherwise required) (2025 FTB 3804‑CR instructions, Part II line 4; 2025 Form 540 booklet, "Line 43 through Line 45"). That is the "Special Credits and Nonrefundable Credits" section: total credits on line 47 are subtracted from line 35 and "If the amount on line 47 is more than the amount on line 35, enter ‑0‑" on line 48 (2025 Form 540 instructions). The refund computation (line 97) compares total payments and refundable credits on line 95 (withholding, 540‑ES payments, Forms 592‑B/593 withholding, CalEITC/YCTC/FYTC, elective refundable film credit) with total tax on line 64; the PTE credit never enters line 95, so it can never generate a refund. Its only relief is the carryover reported on FTB 3804‑CR, Part II, line 2 in later years.

**Ordering and coverage limits that increase the chance of stranded credit.** (i) The PTE credit is applied *after* the other‑state tax credit (§ 17039(a)(6) vs. (a)(7), for TY 2022+), and for OSTC purposes "net tax payable" is grossed up by the PTE credit used (§ 17052.10(f); FTB PTE page). (ii) It may reduce regular tax below tentative minimum tax (§ 17039(c)(1)(AD)) but "cannot reduce the alternative minimum tax" (FTB 3804‑CR instructions). (iii) It cannot offset the 1% Behavioral (Mental) Health Services Tax on taxable income over $1 million, because § 17039's credit‑allowance provisions do not apply to that tax (§ 17043(a), (c)(1); FTB "Help with PTE elective tax"). (iv) The $5 million business‑credit cap does not apply to it (FTB 3804‑CR instructions; § 17039.4 per FTB). (v) It cannot be claimed on a nonresident group return, cannot be assigned, and a nonresident's 7% withholding obligation is unaffected by the election – so a nonresident partner may end up with refundable withholding *and* nonrefundable PTE credit (FTB Help page; FTB 3804‑CR instructions).

**Entity‑level cash flow for 2025.** Two payments: (1) on or before June 15, 2025 (June 16, 2025, next business day), the greater of 50% of the 2024 elective tax paid or $1,000 (§ 19904(a)(2)(A)); (2) the balance by the original due date of the entity's return without extension – March 15, 2026 (March 16, 2026) for a calendar‑year partnership or S corporation (§ 19904(a)(2)(B)). For 2022–2025 years the June 15 payment is a condition of the election: "if no payment is made as required … the qualified entity may not make the election under Section 19900 for that taxable year" (§ 19904(b)); FTB treats an *underpayment* the same way and allows no exceptions absent declared‑disaster relief (FTB Help page). The election itself is made on the timely filed original 2025 return with FTB 3804 and is irrevocable (§ 19900(d)). Payments are made only by Web Pay or FTB 3893 and stay on the entity's account; an entity overpayment is applied to the entity's other liabilities or refunded to the entity – it cannot be moved to owners (FTB PTE page; Help page).

**Partner‑level cash flow.** The FTB does allow the owner to anticipate the credit in estimates: "the PTE elective tax credit does reduce the computation of estimated payments for qualified taxpayers," but not with respect to the 1% behavioral health services tax (FTB Help page). Beyond that, the owner's only liquidity path is using the credit against 2025 "net tax" and carrying any excess to 2026–2030 (§ 17052.10(c)); the five‑year carryover survives the sunset of the credit statute (FTB Help page). No cash refund, no interest, no assignment, and a hard expiry after five carryover years.

## 3. Side‑by‑side (calendar 2025)

| Feature | New York State PTET | California PTE elective tax |
|---|---|---|
| Statutory credit | Tax Law § 606(kkk) | R&TC § 17052.10 |
| Credit amount | Direct share of PTET actually paid/reported by entity (graduated 6.85%–10.9% entity rates); resident/nonresident pools | Flat 9.3% × consenting owner's share of QNI |
| Owner consent | Not required; all direct Art. 22 partners | Required; nonconsenting owners excluded |
| Excess over owner's tax | Treated as overpayment; refunded or credited to next year, without interest (§ 606(kkk)(4)) | Carried forward 5 years, then lost; not refundable, not assignable (§ 17052.10(c); FTB 3804‑CR instr.) |
| Placement on return | Refundable‑credit/payments section: IT‑653 → IT‑201‑ATT line 12 (code 653) → line 18 → IT‑201 line 71 | Nonrefundable special credits: FTB 3804‑CR, code 242, Form 540 lines 43–45/Schedule P; line 48 floored at zero |
| Taxes it can offset | All taxes reported on IT‑201/IT‑203/IT‑205; applied after nonrefundable credits (e.g., § 620 resident credit) | "Net tax" only; after OSTC; can go below TMT but not offset AMT or the 1% § 17043 tax |
| Entity payments | Four quarterly estimates (3/17, 6/16, 9/15, 12/15/2025); balance with annual return 3/16/2026 | June 16, 2025 prepayment (greater of 50% of 2024 tax or $1,000) – missing it forfeits the 2025 election and all credits; balance 3/16/2026 |
| Owner estimates | May take PTET credit into account for IT‑2105 (FAQ); entity payments cannot be transferred to owners | Credit reduces owner's 540‑ES computation (FTB); entity payments cannot be transferred to owners |
| Owner add‑back | § 612(b)(43)(A) add‑back equal to credit claimed | Add‑back is at the entity level in computing California net income (FTB) |
| Timing of claim | Same tax year in which entity files its PTET annual return; amend if filed early | Same taxable year as the entity's election; fiscal‑year entity electing for 2025 → owner claims in 2026 year (§ 17052.10(d)) |

**Illustration.** A partner with a $500,000 distributive share: California credit = 9.3% × $500,000 = $46,500.00; if the partner's 2025 California "net tax" is $40,000, $40,000 is used and $6,500.00 carries to 2026–2030 with no cash back. In New York, if the entity's PTE taxable income is $500,000, PTET = 6.85% × $500,000 = $34,250.00; if the partner's 2025 Article 22 tax is $30,000, the $4,250.00 excess is refunded (without interest) or applied to 2026 estimated tax.

## 4. Caveats

- **2026–2030 California changes (not applicable to 2025).** SB 132 (Stats. 2025, ch. 17) added § 17052.11 and Part 10.4.1 for taxable years beginning on or after January 1, 2026 and before January 1, 2031. The credit remains nonrefundable with the same five‑year carryover (§ 17052.11(c)); a missed or short June 15 payment no longer voids the election but reduces each owner's "qualified amount" by 12.5% of the owner's share of the unpaid amount (§ 17052.11(b)(2)(B); FTB example). § 17052.11 is operative only if the IRC § 164(b)(6) SALT limitation was extended (§ 17052.11(h)); the current text of § 164(b)(6) (as amended by P.L. 119‑21) continues the limitation with an "applicable limitation amount," and FTB's guidance treats the 2026–2030 regime as in effect.
- **New York City PTET.** Form IT‑653 also carries the NYC PTET credit (Tax Law § 1310(g)) on the same refundable line for NYC‑resident partners, with a parallel add‑back under § 612(b)(43‑a).
- **Nonresidents.** Neither state allows the credit on a group nonresident return; California nonresidents remain subject to 7% withholding regardless of the election, and New York nonresidents' credits come from the nonresident credit pool.
- **No interest in New York.** Although refundable, the excess New York credit is refunded "without interest" (§ 606(kkk)(4)), unlike an ordinary overpayment.
- Guidance relied on: NY TSB‑M‑21(1)C,(1)I and the Department's PTET webpage/FAQ (updated 2025–2026) are informational statements of Department policy, not regulations; FTB's PTE webpages are administrative guidance. Statutory text controls where they differ, and none of the above conflicts with the statutes cited.

{"sources": [
{"url": "https://newyork.public.law/laws/n.y._tax_law_section_606", "name": "N.Y. Tax Law § 606(kkk) – Credit for pass-through entity tax (paras. (1)–(5))"},
{"url": "https://newyork.public.law/laws/n.y._tax_law_section_612", "name": "N.Y. Tax Law § 612(b)(3)(C), (b)(43), (b)(43-a) – PTET addback modifications"},
{"url": "https://www.tax.ny.gov/pdf/memos/ptet/m21-1c-1i.pdf", "name": "N.Y. Dept. of Taxation & Finance TSB-M-21(1)C, (1)I, Pass-Through Entity Tax (Aug. 25, 2021)"},
{"url": "https://www.tax.ny.gov/bus/ptet/", "name": "N.Y. Dept. of Taxation & Finance, Pass-through entity tax (PTET) webpage (updated Apr. 3, 2026)"},
{"url": "https://www.tax.ny.gov/bus/ptet/faq.htm", "name": "N.Y. Dept. of Taxation & Finance, Frequently asked questions about the PTET (updated Oct. 3, 2025)"},
{"url": "https://www.tax.ny.gov/pit/credits/pass-through-entity-tax-credit.htm", "name": "N.Y. Dept. of Taxation & Finance, Pass-through entity tax credit webpage"},
{"url": "https://www.tax.ny.gov/pdf/current_forms/it/it653i.pdf", "name": "Form IT-653-I, Instructions for Form IT-653, Pass-Through Entity Tax Credit (2025)"},
{"url": "https://www.tax.ny.gov/pdf/current_forms/it/it653_fill_in.pdf", "name": "2025 Form IT-653, Pass-Through Entity Tax Credit (Tax Law §§ 606(kkk), 1310(g))"},
{"url": "https://www.tax.ny.gov/pdf/current_forms/it/it201att_fill_in.pdf", "name": "2025 Form IT-201-ATT, Part 1, Section D, lines 12–18 (Other refundable credits)"},
{"url": "https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=RTC&sectionNum=17052.10.", "name": "Cal. Rev. & Tax. Code § 17052.10 (PTE elective tax credit, TY 2021–2025; subds. (a)–(k))"},
{"url": "https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=RTC&sectionNum=17052.11.", "name": "Cal. Rev. & Tax. Code § 17052.11 (PTE elective tax credit, TY 2026–2030; SB 132, Stats. 2025, ch. 17)"},
{"url": "https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=RTC&sectionNum=17039.", "name": "Cal. Rev. & Tax. Code § 17039(a)(6)–(8), (c)(1)(AD) – credit ordering and TMT"},
{"url": "https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=RTC&sectionNum=17043.", "name": "Cal. Rev. & Tax. Code § 17043(a), (c)(1) – 1% Mental/Behavioral Health Services Tax not reduced by credits"},
{"url": "https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=RTC&sectionNum=19900.", "name": "Cal. Rev. & Tax. Code § 19900 – PTE elective tax (9.3%, qualified net income, irrevocable election)"},
{"url": "https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=RTC&sectionNum=19904.", "name": "Cal. Rev. & Tax. Code § 19904(a)(2), (b) – June 15 prepayment and balance due; forfeiture of election"},
{"url": "https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=RTC&sectionNum=19906.", "name": "Cal. Rev. & Tax. Code § 19906 – sunset/inoperative provisions for Part 10.4"},
{"url": "https://www.ftb.ca.gov/file/business/credits/pass-through-entity-elective-tax/index.html", "name": "Cal. FTB, Pass-through entity (PTE) elective tax webpage (updated Apr. 23, 2026)"},
{"url": "https://www.ftb.ca.gov/file/business/credits/pass-through-entity-elective-tax/help.html", "name": "Cal. FTB, Help with pass-through entity (PTE) elective tax (credit ordering, estimated taxes, SB 132 changes)"},
{"url": "https://www.ftb.ca.gov/forms/2025/2025-3804-cr-instructions.html", "name": "2025 Instructions for Form FTB 3804-CR, Pass-Through Entity Elective Tax Credit (General Information A, E, F; Part II line 4, credit code 242)"},
{"url": "https://www.ftb.ca.gov/forms/2025/2025-540-booklet.html", "name": "2025 Form 540 Personal Income Tax Booklet – lines 43–48, Payments lines 71–78, lines 97–100, Credit Chart code 242"},
{"url": "https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title26-section164&num=0&edition=prelim", "name": "IRC § 164(b)(6) (as amended by P.L. 119-21) – continuation of SALT deduction limitation"}
]}

91%

Missed checks: the governing New York statute; which California owners are excluded from the credit (disregarded entities).

The question above is a Current & Temporal Analysis task in State & Local Tax: for 2025, the agent must compare how New York’s and California’s pass-through entity tax credits reach an individual partner and whether each state refunds a credit that exceeds the partner’s personal income tax, citing the governing authority in each state. The feedback under each answer summarizes the topics of the rubric checks it missed; most turn on how New York handles credits from more than one electing entity and on supporting details such as the statutory citations, the California claim form, and which owners California excludes from the credit. Missed checks are paraphrased from the rubric. A check fails when the answer does not state the required point, even if it touches the topic.

Timeouts

Each task has a three-hour limit. If the agent has not submitted a final answer when the limit is reached, the run is killed and the question scores 0. Five of the 18 models hit the limit on at least one test question:

ModelTimeouts (of 193)Share of test set
GPT-5.6 Terra189.3%
GPT-5.6 Luna84.1%
GPT-5.6 Sol52.6%
Kimi K342.1%
Qwen 3.8 Max10.5%

The other thirteen models, including all three leaders, answered every test question within the limit. The timeouts are concentrated in the three OpenAI 5.6 checkpoints, which are also the most tool-heavy agents (see Tool Use): their logs end mid-research, typically inside a retrieve_information call dozens of turns in, rather than at a stalled or crashed step. The timed-out questions cluster in Fact-Pattern Analysis (16 of the 36 timeouts across the five models), Current & Temporal Analysis (10), and Forms & Filings (8), the categories that require reading several full documents and reconciling them. For GPT-5.6 Terra the 18 zeros hypothetically cost about seven points: its Accuracy on the 175 questions it did answer is 71.9%, against 65.2% overall, which would move it from tenth to fifth had it finished every task at that rate. GPT-5.6 Luna gains 2.6 points on the same basis (63.4% vs. 60.8%), GPT-5.6 Sol 1.8 (69.8% vs. 68.0%), and Kimi K3 1.4 (70.1% vs. 68.7%).


Methodology

Agents are evaluated on a shared harness with access to five tools. Each task has a three-hour time limit with no fixed step count; the agent may iterate with tools until it submits a final answer or the limit is reached.

  • web_search: searches the web for relevant sources and returns results with titles, URLs, and excerpts from each page
  • lookup_authority: retrieves the text of statutes, regulations, and public laws directly by citation, returning the cited provision along with its canonical source URL
  • fetch_document: downloads a web page or PDF, converts it to plain text, and saves it to the agent’s data storage, allowing the agent to work with documents larger than its context window
  • retrieve_information: queries documents saved in data storage, extracting or summarizing the relevant content without loading the full document into context
  • calculator: evaluates mathematical expressions for exact arithmetic

Tax Agent Bench harness: the agent iterates with search, collect, compute, and recall tools under a three-hour limit, storing fetched documents in a database it can query on demand

Grading

All responses are graded by GPT 5.4 as a judge against each task’s rubric (see Scoring for the weighted metric and must-pass structure). Cited sources are independently verified, and the fraction that resolve to real authoritative documents scales the rubric score.

Dataset

The dataset comprises 391 expert-authored questions, each paired with a gold-standard answer, authoritative sources, and a detailed grading rubric. It is divided into three parts: Public (5 open samples), Private Validation (193 samples available for license), and Test (193 samples).

  • The Public set and agent harness are open and can be accessed here.
  • The Private Validation set is available for license. Interested parties are encouraged to contact us directly for access.
  • The Test set will remain private. Unless otherwise noted, all results reported on this page are based solely on the Test set to prevent overfitting; the only exceptions are the trajectory and sample-answer panels for public question P-004, which are illustrative and do not enter any score.

The Test and Validation splits were sampled to preserve the distribution of question categories.


Citation

If you use this benchmark in your research, please cite:

Citation (BibTeX)

@misc{valsai2026taxagentbench,
title        = {Tax Agent Bench: Evaluating Agents on US Corporate Tax Research},
author       = {Vals AI},
year         = {2026},
month        = sep,
howpublished = {Vals AI},
url          = {https://www.vals.ai/benchmarks/tax_agent_bench},
}