Academic

IOI

Updated 10/6/2026

International Olympiad in Informatics

As of October 6, 2026, Gemini 4 Argon and GPT-6 Astra tie for first on IOI with 100.00%, followed by GPT-6.1 Sol (96.89%).

IOIInformatics olympiad programming problems
ACCURACY

IOI leaderboard

Rank Model Accuracy Cost / Test Input / Output Cost Duration
1 Gemini 4 Argon 100.00% $5.31 $4 / $20 24m19s
2 GPT-6 Astra 100.00% $6.50 $10 / $50 31m40s
3 GPT-6.1 Sol 96.89% $1.26 $2 / $10 31m31s
4 Claude Opus 5.5 95.06% $5.26 $4 / $20 17m15s
5 GPT-5.6 Sol 91.17% $7.98 $4 / $20 1h02m
6 Claude Fable 5.1 90.78% $11.20 $10 / $50 30m06s
7 GPT-5.6 Terra 87.61% $8.66 $2 / $12 1h23m
8 Claude Opus 5 84.33% $16.48 $5 / $25 56m42s
9 Claude Sonnet 5.5 83.06% $7.52 $2 / $10 45m51s
10 GPT-6 Sol 82.61% $2.70 $2 / $10 30m38s
11 Qwen 3.8 Max 68.89% $9.47 $2 / $6 1h47m
12 GLM 5.3 68.44% $7.67 $1.4 / $4.4 1h46m
13 Gemini 3.7 Flash 67.83% $3.40 $1.5 / $7.5 21m09s
14 GPT-5.6 Luna 61.78% $0.58 $0.2 / $1.2 1h10m
15 Hy4 Preview 59.33% $1.35 $0.834 / $2.501 1h01m
16 Grok 4.7 57.72% $12.71 $2 / $6 46m48s
17 Gemini 3.8 Flash 56.94% $3.98 $1.5 / $7.5 14m20s
18 Muse Spark 1.3 Max 56.56% $1.73 $1.25 / $4.25 28m58s
19 GPT-6 Luna 55.56% $0.20 $0.1 / $0.5 35m47s
20 GPT 5.3 Codex 53.83% $2.68 $1.75 / $14 29m52s
21 GLM 5.3 Flash 52.50% $0.43 $0.075 / $0.25 1h32m
22 Gemini 3.1 Pro Preview (02/26) 51.83% $1.25 $2 / $12 12m03s
23 DeepSeek V4 Pro 0813 51.61% $2.15 $1.32 / $3.96 1h07m
24 Kimi K3 48.94% $14.17 $3 / $15 4h18m
25 MiMo V2.6 Flash 47.67% $0.22 $0.14 / $0.28 1h49m
26 Grok 4.6 47.61% $8.18 $2 / $6 3h57m
27 Ember-1 46.33% $24.46 $3 / $15 1h47m
28 Mistral Large 4 45.28% $5.39 $1.36 / $4.18 1h24m
29 Claude Sonnet 5 45.00% $12.89 $2 / $10 59m44s
30 Muse Spark 1.3 43.94% $2.36 $1.25 / $4.25 29m51s
31 Grok 4.5 40.56% $5.64 $2 / $6 1h27m
32 DeepSeek V4.1 Flash 40.28% $0.24 $0.3 / $1.2 18m30s
33 MiMo V2.6 Pro 39.33% $0.84 $0.435 / $0.87 2h26m
34 Qwen 3.8 27B 39.06% $3.55 $0.5 / $3 1h39m
35 Step 5 Preview 38.06% $2.03 $1 / $2.7 1h20m
36 Gemini 3.6 Flash 35.06% $8.67 $1.5 / $7.5 48m23s
37 DeepSeek V4 Flash 0731 32.72% $0.37 $0.44 / $1.32 27m43s
38 Muse Spark 1.2 21.78% $2.64 $1.25 / $4.25 25m35s
39 Inkling 14.94% $1.71 $1 / $4.05 30m12s
40 Inkling Small 9.33% $0.34 $0.3 / $1.2 11m18s
41 Mercury 2.5 2.44% $0.16 $0.2 / $0.75 4m03s

Key Takeaways

  • Unlike the saturated knowledge benchmarks, IOI still sharply separates models: GPT-6 Astra and Gemini 4 Argon solve every problem in all three years, GPT-6.1 Sol (96.89%), Claude Opus 5.5 (95.06%), GPT-5.6 Sol (91.17%), Claude Fable 5.1 (90.78%) and GPT-5.6 Terra (87.61%) follow, and the forty-model field then spreads across nearly a hundred points, down to 2.44%, so competitive-programming ability remains a real differentiator.
  • The agent sees only what a contestant sees: the statement, the sample grader and one sample. It has no tests, no internet and no submission feedback, so every point comes from code the model wrote and checked itself.
  • Scores decline on the newest problems: cohort mean accuracy is 60.78% on 2024 and 56.39% on 2025 but 52.88% on 2026, and twenty-three of the forty models score lowest on the 2026 problems. GPT 5.3 Codex drops from 62.83% on 2024 to 41.67% on 2026. Astra and Terra both score 100% on 2026 even though the contest ran after their training cutoffs, so recency alone does not cap performance.

Why IOI?

Recently, top LLM labs like OpenAI and Google reported that their models achieved gold medals on the International Mathematical Olympiad (IMO). However, advanced models are starting to saturate IMO, meaning it may no longer effectively differentiate between the capabilities of top-performing models. Reports also suggest the evaluation process faced coordination challenges, with AI companies seeking expedited validation mid-competition that may not reflect standard IMO assessment procedures.

The International Olympiad in Informatics (IOI) offers several advantages as an LLM benchmark. Unlike the IMO, the IOI is not yet saturated, providing clear differentiation between model capabilities. The competition features standardized and automated grading, ensuring objective evaluation without subjective scoring. Additionally, the IOI has real-world relevance as it tests C++ programming skills that are directly applicable to software development.


Benchmark Design

We designed our benchmark to imitate competition conditions as closely as possible.

Agent Harness

Each model runs inside OpenCode, a general-purpose coding agent, in an isolated sandbox with a c++ (v20) toolchain. Its workspace holds exactly what the contest hands a contestant: the problem statement, the task header and solution stub, the sample grader or manager, the compile and run scripts, and the sample input and output from the statement. The official test data, the subtask test lists, and the grading script are withheld until the attempt is over. The agent reads, writes, compiles and tests files freely in its workspace, and writes its answer to /workspace/solution.cpp. Like a contestant, it has no access to the public internet: the sandbox reaches only our model gateway, so published editorials and reference solutions are out of reach.

Unlike a contestant, the agent has no submission tool and therefore receives no score feedback during the attempt: it can run the samples and whatever tests it writes for itself, and must decide on its own when its solution is good enough.

Some models spend their entire per-response output budget reasoning about a problem before writing any code. When that happens, the harness continues the cut-off turn on the next step rather than failing the problem: it replays the provider’s own reasoning state where the API returns one, and otherwise restarts the turn with an instruction to work in smaller steps. The same rule applies to every model on this board, and the cost column includes the reasoning spent this way.

Results from our earlier harness, which gave the agent an interactive grading tool, are preserved on IOI v1. The two harnesses are not comparable, so their results are reported on separate boards.

Scoring

Grading matches the olympiad: after the attempt ends, the solution left in the workspace is copied into a fresh grading directory alongside pristine copies of the official test data and grader, compiled, run against every official test, and scored per subtask out of 100 points. Every grading file is verified unchanged after the run. A solution that fails to compile, or that the agent never wrote, scores zero. A model’s score for a year is its mean score across that year’s six problems, and its overall score is the mean of its three yearly scores.


Results

GPT-6 Astra and Gemini 4 Argon reach 100% accuracy, solving all eighteen problems, ahead of GPT-6.1 Sol at 96.89%, Claude Opus 5.5 at 95.06% (every 2024 and 2025 problem solved, 85.17% on 2026), GPT-5.6 Sol at 91.17%, Claude Fable 5.1 at 90.78%, GPT-5.6 Terra at 87.61%, Claude Opus 5 at 84.33% and Claude Sonnet 5.5 at 83.06% (85.00% on 2024, 88.33% on 2025, 75.83% on 2026) and GPT-6 Sol at 82.61%. Fable 5.1 and Opus 5 both solve every 2024 problem and score 87.17% on 2025; Fable holds 85.17% on 2026 while Opus falls to 65.83%. Qwen 3.8 Max (68.89%), GLM 5.3 (68.44%), Gemini 3.7 Flash (67.83%), GPT-5.6 Luna (61.78%), Hy4 Preview (59.33%), Gemini 3.8 Flash (56.94%), Muse Spark 1.3 Max (56.56%), GPT 5.3 Codex (53.83%), GLM 5.3 Flash (52.50%), Gemini 3.1 Pro Preview (02/26) (51.83%), DeepSeek V4 Pro 0813 (51.61%), Kimi K3 (48.94%), MiMo V2.6 Flash (47.67%), Grok 4.6 (47.61%), Ember-1 (46.33%), Claude Sonnet 5 (45.00%) and Muse Spark 1.3 (43.94%) form a middle group, while Grok 4.5 (40.56%), DeepSeek V4.1 Flash (40.28%, one of the cheapest models on the board at $0.24 per problem), MiMo V2.6 Pro (39.33%), Qwen 3.8 27B (39.06%, thirty points behind Qwen 3.8 Max), Step 5 Preview (38.06%; 26.67% on 2024 but 49.83% on 2025), Gemini 3.6 Flash (35.06%), DeepSeek V4 Flash 0731 (32.72%), Muse Spark 1.2 (21.78%), Inkling (14.94%) and Inkling Small (9.33%) trail the field. Both Muse Spark 1.3 variants at least double the score of their predecessor Muse Spark 1.2. Claude Sonnet 5 is the fourth most expensive model on the board at $12.89 per problem, behind only Ember-1, Opus 5 and Kimi K3: on thirteen of its eighteen problems it spent its entire 128k-token output budget reasoning before writing any code, and the harness had to continue the cut-off turn.

We evaluate three years to check for data contamination. The 2026 problems were released only after the models in this cohort were trained, and they are the hardest set for twenty-three of the forty: GPT 5.3 Codex scores 62.83% on 2024 and 57.00% on 2025 but 41.67% on 2026, and cohort mean accuracy falls from 60.78% (2024) and 56.39% (2025) to 52.88% (2026). GPT-6 Astra, Gemini 4 Argon and GPT-5.6 Terra all score 100% on 2026, so the newest set is not out of reach.