Proprietary

MedCode

Updated 10/1/2026

Can models support the medical billing process?

As of October 1, 2026, Claude Opus 5 ranks first on MedCode with 63.57%, followed by Gemini 3.1 Pro Preview (02/26) at 59.06% and Gemini 4 Argon (58.80%).

MedCodeMedical billing code support
ACCURACY

MedCode leaderboard

Rank Model Accuracy Cost In / Out Latency
1 Claude Opus 5 63.57% $5 / $25 24.50s
2 Gemini 3.1 Pro Preview (02/26) 59.06% $2 / $12 38.52s
3 Gemini 4 Argon 58.80% $4 / $20 79.53s
4 Claude Fable 5 56.07% $10 / $50 91.44s
5 Gemini 3 Flash (12/25) 55.92% $0.5 / $3 44.15s
6 Gemini 3.5 Flash 55.83% $1.5 / $9 25.29s
7 Claude Opus 4.7 54.86% $5 / $25 54.25s
8 Claude Fable 5.1 53.51% $10 / $50 3m35s
9 Gemini 3.7 Flash 53.39% $1.5 / $7.5 9.16s
10 Claude Opus 4.8 53.22% $5 / $25 105.92s
11 Gemini 3.6 Flash 53.15% $1.5 / $7.5 16.68s
12 Claude Sonnet 5.5 52.92% $2 / $10 3m59s
13 GPT 5.1 52.73% $1.25 / $10 54.55s
14 Gemini 3 Pro (11/25) 52.20% $2 / $12 55.84s
15 Muse Spark 51.31% N/A 2m03s
16 Gemini 2.5 Pro 50.59% $1.25 / $10 27.41s
17 Claude Opus 5.5 49.80% $4 / $20 4m07s
18 GPT 5.2 49.75% $1.75 / $14 2m37s
19 GPT 5 49.63% $1.25 / $10 58.27s
20 Grok 4.7 49.55% $2 / $6 3m33s
21 Kimi K3 49.36% $3 / $15 36.99s
22 Muse Spark 1.2 49.35% $1.25 / $4.25 60.17s
23 Claude Opus 4.5 (Thinking) 49.16% $5 / $25 60.84s
24 Claude Opus 4.6 (Thinking) 49.13% $5 / $25 2m36s
25 GPT 5.5 49.10% $5 / $30 2m40s
26 GPT-6.1 Sol 48.84% $2 / $10 115.34s
27 GPT-6 Astra 48.49% $10 / $50 102.86s
28 Claude Opus 4.6 (Nonthinking) 48.24% $5 / $25 4.06s
29 Gemini 3.8 Flash 48.13% $1.5 / $7.5 42.89s
30 Gemini 3.1 Flash Lite Preview 47.60% $0.25 / $1.5 8.55s
31 Claude Sonnet 5 47.54% $2 / $10 2m15s
32 o3 47.29% $2 / $8 17.68s
33 Claude Opus 4.1 (Thinking) 47.23% $15 / $75 33.26s
34 GPT-6 Sol 47.07% $2 / $10 62.71s
35 MiniMax-M3 46.29% $0.6 / $2.4 63.12s
36 Claude Opus 4.5 (Nonthinking) 45.17% $5 / $25 5.02s
37 MiMo V2.6 Pro 44.97% $0.435 / $0.87 2m57s
38 Grok 4.6 44.71% $2 / $6 2m07s
39 GPT-6 Luna 44.69% $0.1 / $0.5 108.77s
40 Claude Sonnet 4.5 (Thinking) 44.13% $3 / $15 74.33s
41 GPT-5.6 Sol 43.97% $4 / $20 96.59s
42 Gemini 3.5 Flash Lite 43.49% $0.3 / $2.5 6.02s
43 GPT-5.6 Terra 43.41% $2 / $12 18.41s
44 Grok 4.5 43.29% $2 / $6 59.64s
45 Hy4 Preview 43.25% $0.834 / $2.501 6m39s
46 GPT 5 Mini 43.05% $0.25 / $2 28.17s
47 GLM 5.3 42.86% $1.4 / $4.4 2m49s
48 DeepSeek V4 Pro 0813 42.47% $1.32 / $3.96 3m01s
49 GPT-5.6 Luna 42.39% $0.2 / $1.2 81.28s
50 GLM 5.1 41.60% $1 / $3.2 77.58s
51 DeepSeek V4 Flash 0731 41.41% $0.44 / $1.32 2m07s
52 Claude Opus 4.1 (Nonthinking) 41.37% $15 / $75 13.08s
53 GPT 5.4 (xhigh) 41.29% $2.5 / $15 3m07s
54 Inkling 41.19% $1 / $4.05 2m47s
55 DeepSeek V4.1 Flash 41.17% $0.3 / $1.2 35.13s
56 MiMo V2.6 Flash 41.06% $0.14 / $0.28 93.37s
57 GPT 5.4 Nano 41.03% $0.2 / $1.25 8.04s
58 GLM 5.2 40.77% $1.4 / $4.4 95.27s
59 Qwen 3.8 Max 40.67% $2 / $6 6m18s
60 Claude Sonnet 4.5 (Nonthinking) 40.57% $3 / $15 12.01s
61 Gemini 2.5 Flash Preview (9/25) (Nonthinking) 40.54% $0.3 / $2.5 12.70s
62 DeepSeek V4 40.45% $1.32 / $3.96 6m23s
63 Gemini 2.5 Flash (7/17) (Thinking) 40.36% $0.3 / $2.5 22.83s
64 Gemini 2.5 Flash Preview (9/25) (Thinking) 40.33% $0.3 / $2.5 16.11s
65 Kimi K2.6 40.14% $0.95 / $4 5m05s
66 Kimi K2.5 39.32% $0.6 / $3 75.45s
67 Qwen 3.7 Max 38.75% $2.5 / $7.5 28.75s
68 Nemotron 3 Ultra 38.62% N/A 18.68s
69 Gemini 2.5 Flash (7/17) (Nonthinking) 38.42% $0.3 / $2.5 22.14s
70 Grok 4 38.08% $3 / $15 89.00s
71 Grok 4.3 38.07% $1.25 / $2.5 43.92s
72 Inkling Small 37.89% $0.3 / $1.2 3m11s
73 Grok 4 Fast (Reasoning) 37.38% $0.2 / $0.5 17.52s
74 Qwen 3.6 Plus 36.89% $0.5 / $3 56.22s
75 Llama 4 Maverick 36.51% $0.22 / $0.88 21.24s
76 Claude Sonnet 4 (Thinking) 34.96% $3 / $15 39.80s
77 MiniMax-M2.7 34.44% $0.3 / $1.2 30.25s
78 Gemini 2.5 Flash Lite (9/25) (Thinking) 34.19% $0.1 / $0.4 10.49s
79 MiniMax-M2.1 34.08% $0.3 / $1.2 19.55s
80 Claude Sonnet 4 (Nonthinking) 33.94% $3 / $15 7.30s
81 o4 Mini 33.79% $1.1 / $4.4 21.06s
82 Mistral Medium 3.5 33.75% $1.5 / $7.5 34.88s
83 Qwen 3.5 Flash 33.00% $0.1 / $0.4 63.07s
84 GLM 4.7 32.77% $0.6 / $2.2 2m03s
85 Claude Haiku 4.5 (Thinking) 32.68% $1 / $5 34.29s
86 MiMo V2.5 Pro 32.48% $0.435 / $0.87 38.45s
87 Ling 3.0 Flash 32.27% $0.075 / $0.22 9.65s
88 Grok 4.20 (Reasoning) 32.16% $2 / $6 16.55s
89 MiMo V2.5 31.89% $0.14 / $0.28 17.70s
90 Qwen 3 VL Plus 31.65% $0.2 / $1.6 9.95s
91 Qwen 3 Max Thinking 31.37% $1.2 / $6 3m02s
92 Mercury 2.5 31.33% $0.2 / $0.75 7.93s
93 GPT 5 Nano 30.44% $0.05 / $0.4 29.74s
94 Grok 4 Fast (Non-Reasoning) 30.04% $0.2 / $0.5 18.21s
95 Ling 3.0 Flash Fin 29.30% $0.06 / $0.18 31.17s
96 Qwen 3.8 27B 28.70% $0.5 / $3 89.38s
97 Grok 4.1 Fast Non-Reasoning 28.35% $0.2 / $0.5 4.50s
98 Grok 4.1 Fast (Reasoning) 28.08% $0.2 / $0.5 46.50s
99 Gemini 2.5 Flash Lite (Nonthinking) 27.11% $0.1 / $0.4 5.75s
100 Gemini 2.5 Flash Lite (9/25) (Nonthinking) 27.08% $0.1 / $0.4 6.24s
101 Llama 4 Scout 23.31% $0.18 / $0.59 10.74s
102 Laguna M.1 23.11% N/A 71.00s
103 Laguna XS.2 21.25% N/A 33.09s
104 Command A+ 19.72% N/A 103.75s

Partners in Evaluation

Key Takeaways

  • AI models are being deployed to automate medical coding, but current systems still struggle with accuracy — the leading model reaches only 63.57% on this critical healthcare task.
  • Claude Opus 5 leads overall at 63.57%, ahead of Gemini 3.1 Pro Preview (02/26) (59.06%) and Claude Fable 5 (56.07%). It also leads on secondary-code accuracy (65.36%) and full subcategory precision (64.92%).
  • Models perform better on physical conditions like diabetes and hypertension, but struggle significantly with mental health diagnoses and other complex categories, raising concerns about reliability in real-world clinical settings.

Results

Claude Opus 5 takes the top spot at 63.57%, ahead of Gemini 3.1 Pro Preview (02/26) (59.06%), Claude Fable 5 (56.07%), Gemini 3 Flash (12/25) (55.92%), and Gemini 3.5 Flash (55.83%).

MedCode Performance

Background

The process of accurate medical coding is extraordinarily complex. There are tens of thousands of diagnosis codes, each with detailed rules dictating sequencing, modifiers, and payer-specific requirements1. These rules vary across patient demographics, medical providers, and geographical jurisdictions, making coding extremely difficult. Any error, even those that are seemingly minor, can lead to claim denials and lost revenue for the hospital2. While AI coding systems are being adopted to improve both accuracy and finances, there is little to no oversight or evaluation for how these systems perform under real-world conditions3.

Through our MedCode benchmark, we evaluate AI systems’ ability to perform medical coding given realistic documentation constraints. In collaboration with Protege, we created a dataset of 2755 primary and secondary diagnosis codes for each de-identified patient record, enabling coding evaluation for an entire hospitalization stay from admission through discharge.


Methodology

The International Classification of Diseases (ICD) is the global standard for classifying diseases, symptoms, and injuries as it enables consistent reporting across health systems. We evaluate AI models on the CMS ICD-10-CM standard, which is the official coding standard mandated by the Centers for Medicare & Medicaid services (CMS), as it is publicly available with comprehensive, up-to-date documentation.

Our dataset consists of de-identified patient records containing discharge summaries supported by corresponding progress notes and consult notes throughout the duration of the patient’s hospital stay. We matched the three-digit ZIP code of each patient’s stay to geographical jurisdictions and other demographic information through reference tables, allowing us to obtain key information necessary for accurate coding without compromising patient privacy. Each sample was independently annotated by two certified professional coders, who assigned primary and secondary diagnosis codes based on the ICD-10-CM standard. They were permitted to use Codify, a professional tool for code lookup and validation, but were not allowed to use AI models. Differences between coder annotations were resolved collaboratively to ensure consistency and high quality.

Each model was given the de-identified patient records and prompted to output corresponding primary and secondary diagnoses following the ICD-10-CM standard. All models are evaluated with temperature 1, and produce at most 30k tokens.


Additional Findings

We break down model performance across several dimensions: primary versus secondary diagnoses, the most frequently occurring codes, ICD-10 chapters, and levels of code precision.

Primary vs Secondary Pass Rates

Nearly all models perform better at predicting primary diagnosis codes than secondary codes; Opus 5 and Claude Opus 5.5 are the only exceptions. Interestingly, very few errors stem from misclassifying a code as primary versus secondary. The majority of errors occur when models completely miss a diagnosis code that appears in the rubric.

Primary vs Secondary Pass Rates
CorrectMismatchMissing

Most Frequent Codes

No model consistently leads across the top 10 most occurring diagnoses codes. This inconsistency suggests brittle generalization rather than robust clinical understanding.

Models generally perform well on physiological conditions such as type 2 diabetes (E11.9), hypertension (I10), and acute kidney failure (N17.9), but they struggle with mental-health diagnoses like F32.A (depression). Notably, Opus 5, Claude Fable 5.1, and GPT 5.1 each correctly identify F32.A in 80% of cases.

Most Frequent ICD-10 Codes Pass Rates by Model

Pass Rates by Chapter

This pattern extends to the chapter level. Models perform better on codes tied to physical conditions, while accuracy drops sharply for chapters like F (Mental, Behavioral, and Neurodevelopmental disorders) and Z (Factors influencing health status and contact with health services).

O (delivery) codes were excluded because obfuscation in the dataset made them unreliable to evaluate.

ICD-10 Chapter Pass Rates by Model

Pass Rates by Code Precision

We also investigated how models perform when diagnosis codes are evaluated with greater precision. The plot below shows pass rates at different code levels. For example, for the code F32.A, the chapter is F (mental disorders), the category is F32 (depression), and the subcategory is the full code F32.A (major depressive disorder, single episode, psychotic features).

The accuracy drop-off with increased precision is expected. However, even when evaluating by category alone, the best models still perform poorly, indicating that the primary cause of failure is completely missing the diagnosis, not mislabelling the specificity.

Precision Pass Rates by Model

Citations

[1] ICD.Codes. (2018, January 19). ICD-9 to ICD-10 explained. Retrieved from https://web.archive.org/web/20180119120210/https://icd.codes/articles/icd9-to-icd10-explained

[2] BlueBrix Health. (2025, March 12). The hidden costs of coding errors: How accurate medical coding boosts revenue. BlueBrix Health. https://bluebrix.health/blogs/the-hidden-costs-of-coding-errors-how-accurate-medical-coding-boosts-revenue

[3] Angus, D. C., Lee, S. I., Beam, A. L., Denecke, K., Farahani, N. I., Kohane, I. S., Matheny, M. E., Sendak, M. P., Shah, N. H., Steinhubl, S. R., & Topol, E. J. (2025). AI, health, and health care today and tomorrow: The JAMA summit report on artificial intelligence. JAMA, 333(14), 1279–1287. https://doi.org/10.1001/jama.2025.18490