Archived Benchmark

Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases. Previous results are preserved here for posterity.

Proprietary

TaxEval

Updated 9/1/2026

A Vals-created set of questions and responses to tax questions

As of September 1, 2026, Muse Spark 1.2 ranks first on TaxEval with 80.38%, followed by Muse Spark 1.1 (79.72%) and Muse Spark (77.68%).

TaxEval v2Tax question answering accuracy
ACCURACY

TaxEval leaderboard

Rank Model Accuracy Cost In / Out Latency
1 Muse Spark 1.2 80.38% $1.25 / $4.25 48.31s
2 Muse Spark 1.1 79.72% $1.25 / $4.25 32.82s
3 Muse Spark 77.68% N/A 57.52s
4 Claude Sonnet 4.6 77.11% $3 / $15 2m08s
5 Claude Fable 5 76.94% $10 / $50 56.87s
6 GPT-5.6 Terra 76.17% $2 / $12 37.28s
7 GPT-5.6 Luna 76.17% $0.2 / $1.2 76.52s
8 Claude Opus 4.6 (Thinking) 75.96% $5 / $25 79.25s
9 Claude Fable 5.1 75.96% $10 / $50 85.56s
10 Grok 3 75.88% $3 / $15 13.06s
11 GPT 5.2 75.76% $1.75 / $14 60.48s
12 Kimi K3 75.72% $3 / $15 2m25s
13 Grok 4 Fast (Reasoning) 75.70% $0.2 / $0.5 10.18s
14 Claude Opus 4.8 75.63% $5 / $25 70.33s
15 Claude Sonnet 5 75.63% $2 / $10 3m22s
16 GLM 5.3 Flash 75.59% $0.075 / $0.25 53.06s
17 Qwen 3.8 Max 75.55% $2 / $6 4m26s
18 Inkling Small 75.51% $0.3 / $1.2 4m02s
19 Qwen 3.7 Max 75.31% $2.5 / $7.5 57.68s
20 Inkling 75.31% $1 / $4.05 3m18s
21 Claude Opus 4.7 75.27% $5 / $25 62.07s
22 GPT 5 Mini 75.22% $0.25 / $2 37.43s
23 Claude Opus 5 75.14% $5 / $25 58.34s
24 GPT 4.1 75.06% $2 / $8 7.89s
25 GPT 5.5 74.98% $5 / $30 95.83s
26 Gemini 3.6 Flash 74.86% $1.5 / $7.5 12.34s
27 GPT 5.1 74.86% $1.25 / $10 44.18s
28 Claude Opus 4.5 (Thinking) 74.86% $5 / $25 47.45s
29 o4 Mini 74.78% $1.1 / $4.4 15.45s
30 GPT-5.6 Sol 74.78% $4 / $20 62.03s
31 Gemini 3.7 Flash 74.73% $1.5 / $7.5 7.83s
32 Qwen 3.6 Plus 74.73% $0.5 / $3 70.63s
33 Kimi K2.6 74.65% $0.95 / $4 4m27s
34 o3 74.57% $2 / $8 24.02s
35 GPT 4o (2024-11-20) 74.53% $2.5 / $10 5.86s
36 Gemini 3.8 Flash 74.45% $1.5 / $7.5 8.76s
37 Gemini 3.5 Flash 74.37% $1.5 / $9 19.69s
38 Claude Opus 4.5 (Nonthinking) 74.33% $5 / $25 14.88s
39 o1 74.28% $15 / $60 19.48s
40 Kimi K2.5 74.20% $0.6 / $3 90.46s
41 Grok 4.20 (Reasoning) 74.12% $2 / $6 13.39s
42 Claude 3.7 Sonnet (Thinking) 74.04% $3 / $15 43.31s
43 Qwen 3 Max Preview 73.96% $1.2 / $6 21.41s
44 GPT 5.4 (xhigh) 73.96% $2.5 / $15 78.72s
45 Gemini 3 Flash (12/25) 73.88% $0.5 / $3 12.00s
46 MiMo V2.5 Pro 73.79% $0.435 / $0.87 62.83s
47 Claude Opus 4.1 (Thinking) 73.67% $15 / $75 26.64s
48 Qwen 3 Max 73.51% $1.2 / $6 29.30s
49 GPT 5 73.39% $1.25 / $10 78.67s
50 GLM 5.2 73.34% $1.4 / $4.4 64.94s
51 Claude Sonnet 4.5 (Thinking) 73.30% $3 / $15 48.39s
52 Grok 4.1 Fast (Reasoning) 73.14% $0.2 / $0.5 30.76s
53 Nemotron 3 Ultra 73.10% N/A 36.48s
54 Mistral Large 3 73.06% $0.5 / $1.5 29.53s
55 DeepSeek V4 Pro 0813 73.06% $1.32 / $3.96 3m46s
56 Grok 3 Mini Reasoning 72.98% $0.3 / $0.5 11.73s
57 Gemini 2.5 Pro Exp 72.89% $1.25 / $10 21.77s
58 Gemini 3.1 Pro Preview (02/26) 72.88% $2 / $12 41.93s
59 Gemini 2.5 Flash Preview (9/25) (Nonthinking) 72.73% $0.3 / $2.5 6.27s
60 MiniMax-M3 72.73% $0.6 / $2.4 94.04s
61 Gemini 3.5 Flash Lite 72.61% $0.3 / $2.5 8.81s
62 Gemini 3 Pro (11/25) 72.57% $2 / $12 24.40s
63 Claude 3.7 Sonnet (Nonthinking) 72.40% $3 / $15 6.28s
64 Gemini 2.5 Flash Preview (9/25) (Thinking) 72.40% $0.3 / $2.5 10.63s
65 GLM 4.5 72.40% $0.6 / $2.2 2m09s
66 GLM 5.3 72.36% $1.4 / $4.4 98.32s
67 DeepSeek R1 72.28% $1.35 / $5.4 2m33s
68 Qwen 3.5 Flash 72.16% $0.1 / $0.4 73.00s
69 DeepSeek V4 72.08% $1.32 / $3.96 3m10s
70 Claude Sonnet 4 (Thinking) 72.00% $3 / $15 32.38s
71 GPT 4.1 Mini 71.91% $0.4 / $1.6 5.92s
72 Claude Opus 4 (Nonthinking) 71.91% $15 / $75 15.23s
73 MiMo V2.5 71.83% $0.14 / $0.28 26.30s
74 Gemini 3.1 Flash Lite Preview 71.79% $0.25 / $1.5 7.64s
75 Kimi K2 Thinking 71.71% $0.6 / $2.5 61.41s
76 Grok 4.5 71.67% $2 / $6 3m15s
77 GPT OSS 120B 71.59% $0.15 / $0.6 47.66s
78 Grok 4 Fast (Non-Reasoning) 71.58% $0.2 / $0.5 4.25s
79 Claude Opus 4.1 (Nonthinking) 71.46% $15 / $75 32.60s
80 Qwen 3.6 27B 71.26% $0.6 / $3.6 2m28s
81 GPT 5.4 Mini 71.22% $0.75 / $4.5 44.58s
82 GLM 5.1 71.19% $1 / $3.2 84.07s
83 Gemini 2.5 Flash Preview 4/17 (Nonthinking) 71.18% $0.3 / $2.5 5.39s
84 GPT 4o (2024-08-06) 71.14% $2.5 / $10 9.54s
85 Grok 3 Mini Reasoning 71.14% $0.3 / $0.5 25.45s
86 DeepSeek V3 (03/24/2025) 71.10% $0.9 / $0.9 34.01s
87 Grok 4.6 71.10% $2 / $6 47.57s
88 Qwen 3.8 27B 70.85% $0.5 / $3 94.49s
89 Grok 4.3 70.81% $1.25 / $2.5 6m32s
90 DeepSeek V4 Flash 0731 70.69% $0.44 / $1.32 72.20s
91 Ling 3.0 Flash 70.65% $0.075 / $0.22 7.54s
92 Qwen 3 (235B) 70.65% $0.22 / $0.88 92.74s
93 Gemini 2.5 Flash Preview 4/17 (Thinking) 70.52% $0.3 / $2.5 15.36s
94 Mistral Medium 3.1 (05/2025) 70.32% $0.4 / $2 11.05s
95 Kimi K2 Instruct 70.20% $1 / $3 12.83s
96 Claude 3.5 Sonnet Latest 70.16% $3 / $15 5.94s
97 GLM 5 70.03% $1 / $3.2 2m29s
98 Gemini 2.0 Flash Thinking Exp 69.79% $0.1 / $0.7 12.81s
99 Claude Sonnet 4 (Nonthinking) 69.62% $3 / $15 11.85s
100 o3 Mini 69.42% $1.1 / $4.4 74.31s
101 GLM 4.7 68.77% $0.6 / $2.2 3m07s
102 Gemini 2.0 Pro Exp 68.15% $1.25 / $5 8.99s
103 MiniMax-M2.5 68.15% $0.3 / $1.2 73.24s
104 DeepSeek V3.2 (Thinking) 68.15% $0.56 / $1.68 98.92s
105 Mistral Medium 3.5 67.99% $1.5 / $7.5 102.39s
106 DeepSeek V3 67.91% $0.9 / $0.9 27.70s
107 Gemini 2.0 Flash Exp 67.74% $0.075 / $0.3 7.49s
108 Claude Haiku 4.5 (Thinking) 67.54% $1 / $5 47.52s
109 GPT 5.4 Nano 67.42% $0.2 / $1.25 15.90s
110 GPT 5 Nano 67.38% $0.05 / $0.4 69.94s
111 Grok 2 67.05% $2 / $10 10.62s
112 Llama 4 Maverick 66.56% $0.22 / $0.88 25.14s
113 MiniMax-M2.7 66.56% $0.3 / $1.2 28.69s
114 MiniMax-M2.1 66.35% $0.3 / $1.2 113.26s
115 GLM 4.6 66.23% $0.6 / $2.2 64.15s
116 Gemini 2.5 Flash Lite (9/25) (Nonthinking) 66.23% $0.1 / $0.4 3.60s
117 Gemini 2.0 Flash (001) 65.25% $0.1 / $0.4 5.70s
118 Grok 4 65.09% $3 / $15 2m39s
119 Gemini 2.5 Flash Lite (9/25) (Thinking) 64.72% $0.1 / $0.4 3.94s
120 Mistral Large (11/2024) 63.78% $2 / $6 12.02s
121 Llama 3.3 Nemotron Super (Thinking) 63.74% N/A 32.67s
122 GPT OSS 20B 63.70% $0.07 / $0.3 49.00s
123 Magistral Medium 1.2 (09/2025) 61.94% $2 / $5 46.25s
124 Grok 4.1 Fast Non-Reasoning 61.86% $0.2 / $0.5 9.28s
125 Command A 61.37% $2.5 / $10 11.21s
126 Jamba 1.6 Large 60.88% $2 / $8 16.44s
127 Llama 3.1 Instruct Turbo (405B) 60.88% $3.5 / $3.5 23.55s
128 GPT 4.1 Nano 60.75% $0.1 / $0.4 3.00s
129 GPT 4o Mini 60.55% $0.15 / $0.6 8.90s
130 Magistral Small 1.2 (09/2025) 60.30% $0.5 / $1.5 16.44s
131 Llama 3.3 Nemotron Super (Nonthinking) 60.22% N/A 18.31s
132 Gemini 1.5 Pro (002) 59.48% $1.25 / $5 9.86s
133 Llama 3.3 Instruct Turbo (70B) 59.44% $0.88 / $0.88 3.84s
134 Laguna XS.2 58.95% N/A 27.45s
135 Mistral Small 3.1 (03/2025) 58.30% $0.075 / $0.3 7.95s
136 Jamba 1.5 Large 58.18% $2 / $8 20.03s
137 Claude 3.5 Haiku Latest 57.36% $0.8 / $4 4.99s
138 Llama 3.1 Instruct Turbo (70B) 56.17% $0.88 / $0.88 4.33s
139 Llama 4 Scout 55.19% $0.18 / $0.59 7.04s
140 Mistral Small (02/2024) 49.14% $0.2 / $0.6 9.02s
141 Gemini 1.5 Flash (002) 48.20% $0.075 / $0.3 3.70s
142 Jamba 1.6 Mini 44.60% $0.2 / $0.4 4.34s
143 Jamba 1.5 Mini 41.86% $0.2 / $0.4 4.78s
144 Llama 3.1 Instruct Turbo (8B) 32.34% $0.18 / $0.18 2.46s
145 Laguna M.1 1.64% N/A 3m33s

Key Takeaways

  • Model choice barely matters at the top: the leaders top out around 80% and the top ten sit within about five points of each other, so picking among frontier models buys little accuracy here.
  • The benchmark’s two evaluation dimensions expose separate ceilings: Muse Spark 1.2 scores 92.72% on stepwise reasoning but only 68.03% on answer correctness. The nearly 25-point gap shows that despite strong reasoning demonstration, the frontier challenge is accurate final computation, not reasoning structure.

Dataset and Context

TaxEval v2 evaluates models’ abilities to answer hard tax-related questions. This version focuses on both answer correctness and structured reasoning capabilities. This dataset was created in collaboration with financial and tax experts, who have both created and double-checked all questions and answers.

Some key features:

  • 1,500+ total questions across validation and test sets
  • A balanced distribution of topics and question types
  • Comprehensive evaluation of both answers and reasoning steps

The benchmark consists of three main components:

  • Public Validation Set: 20 samples, available upon request (contact contact@vals.ai).
  • Private Validation Set: 300 samples for model evaluation, available for purchase to evaluate and improve models.
  • Test Set: 1,223 samples. These samples are never shared.

The benchmark is composed of two tasks (each task uses the same questions, but they are evaluated differently):

  1. Answer Correctness: The factual correctness of the answer, as compared to a ground truth.
  2. Stepwise Reasoning: The quality and structure of the reasoning process as compared to the reasoning process of human experts. This ensures models not only provide correct answers but also demonstrate a clear thinking process.

The benchmark includes a diverse range of question types:

  • Application and Compliance (18.3%)
  • Comparative Analysis (16.2%)
  • Numerical Reasoning (16.7%)
  • Problem Solving and Critical Thinking (16.5%)
  • Semantic Analysis (18.0%)
  • Updates and Current Affairs (15.9%)

Each category is carefully balanced between the private validation and test sets to ensure representative sampling. The evaluation process uses Claude Sonnet 4.5 (Nonthinking) as judge for both answer correctness and stepwise reasoning assessment.


Results



Model Output Examples

The hardest questions for the models typically require complex, multi-step calculations, reasoning on what laws and numbers to use given the information, or accessing more recent data.

Below is an example of a question where one model arrives at the correct answer ($18,280) while the others perform the calculations incorrectly, despite all models using a significant amount of reasoning tokens.

Question

Michael, a married taxpayer filing jointly, has an adjusted gross income (AGI) of $300,000 for 2023, excluding any investment income. During the year, he received $20,000 in interest from corporate bonds and $15,000 in interest from municipal bonds issued by his state of residence. In addition, he sold a collectible artwork for $100,000 that he had purchased 5 years ago for $60,000. Calculate Michael's tax liability related to his investment income, including any applicable taxes on capital gains and considering the Net Investment Income Tax (NIIT).

Michael owes a total of $18,280 in taxes related to his investment income.

CORRECT


Changelog

This benchmark was updated on 11/17/25 to switch the LLM-as-judge from Claude 3.5 Sonnet, which had been deprecated and is no longer usable, to Claude 4.5 Sonnet.