Archived Benchmark

Terminal-Bench 2.1 has been superseded by Terminal-Bench 4.0, which replaces it in the Coding sector of Vals Index v2.1. We no longer run this benchmark on new model releases. Previous results are preserved here for posterity.

Academic

Terminal-Bench 2.1

Updated 9/28/2026

State-of-the-art set of difficult terminal-based tasks

As of September 28, 2026, Claude Opus 5.5 (Terminus 2) ranks first on Terminal-Bench 2.1 with 87.64%, followed by GPT-6 Astra (Terminus 2) at 87.27% and GPT-5.6 Sol (Terminus 2) at 85.77%.

Terminal-Bench 2.1Difficult terminal-based coding tasks
ACCURACY

Terminal-Bench 2.1 leaderboard

Rank Model Harness Score Cost / Test Input / Output Cost Duration
1 Claude Opus 5.5 Terminus 2 87.64% $0.50 $4 / $20 6m23s
2 GPT-6 Astra Terminus 2 87.27% $1.34 $10 / $50 4m34s
3 GPT-5.6 Sol Terminus 2 85.77% $1.02 $4 / $20 6m13s
4 Claude Fable 5.1 Terminus 2 85.02% $3.07 $10 / $50 8m20s
5 Claude Opus 5 Terminus 2 84.64% $0.89 $5 / $25 12m33s
6 GPT-6 Sol Terminus 2 83.15% $0.37 $2 / $10 4m59s
7 Claude Sonnet 5.5 Terminus 2 83.15% $0.63 $2 / $10 7m46s
8 Gemini 3.8 Flash Terminus 2 81.27% $1.54 $1.5 / $7.5 9m15s
9 Kimi K3 Terminus 2 80.90% $0.34 $3 / $15 10m09s
10 Claude Fable 5 Terminus 2 80.52% $1.43 $10 / $50 8m25s
11 GPT-5.6 Luna Terminus 2 79.03% $0.05 $0.2 / $1.2 5m56s
12 Muse Spark 1.3 Max Terminus 2 79.03% $0.60 $1.25 / $4.25 12m12s
13 Grok 4.6 Terminus 2 78.28% $0.45 $2 / $6 10m43s
14 Gemini 3.7 Flash Terminus 2 77.53% $1.02 $1.5 / $7.5 8m12s
15 GPT-5.6 Terra Terminus 2 77.53% $0.47 $2 / $12 8m40s
16 GPT 5.5 Terminus 2 76.40% $0.74 $5 / $30 7m07s
17 MiMo V2.6 Flash Terminus 2 76.40% $0.02 $0.14 / $0.28 13m30s
18 DeepSeek V4.1 Flash Terminus 2 74.53% $0.10 $0.3 / $1.2 10m05s
19 Claude Sonnet 5 Terminus 2 74.53% $0.53 $2 / $10 10m35s
20 Gemini 3.5 Flash Terminus 2 74.16% $0.87 $1.5 / $9 6m13s
21 Gemini 3.6 Flash Terminus 2 73.78% $1.17 $1.5 / $7.5 9m12s
22 Grok 4.7 Terminus 2 73.41% $1.08 $2 / $6 13m18s
23 GPT-6 Luna Terminus 2 73.03% $0.03 $0.1 / $0.5 7m05s
24 Muse Spark 1.3 Terminus 2 72.28% $0.66 $1.25 / $4.25 14m18s
25 Claude Opus 4.8 Terminus 2 71.91% $2.41 $5 / $25 15m30s
26 GLM 5.3 Terminus 2 71.54% $0.31 $1.4 / $4.4 13m07s
27 Gemini 3.1 Pro Preview (02/26) Terminus 2 70.79% $0.58 $2 / $12 6m50s
28 Claude Opus 4.8 Claude Code 69.66% $0.97 N/A 6m00s
29 Muse Spark 1.2 Terminus 2 69.66% $0.50 $1.25 / $4.25 14m39s
30 Muse Spark 1.1 Terminus 2 69.29% $0.42 $1.25 / $4.25 10m25s
31 Claude Opus 4.7 Terminus 2 68.54% $1.98 $5 / $25 13m37s
32 Grok 4.5 Terminus 2 67.79% $0.28 $2 / $6 10m24s
33 GLM 5.2 Terminus 2 67.79% $0.43 $1.4 / $4.4 13m41s
34 MiMo V2.6 Pro Terminus 2 67.79% $0.07 $0.435 / $0.87 15m39s
35 Qwen 3.8 Max Terminus 2 67.42% $0.45 $2 / $6 14m51s
36 DeepSeek V4 Flash 0731 Terminus 2 67.04% $0.02 $0.44 / $1.32 11m23s
37 Kimi K2.7 Code Terminus 2 67.04% $0.26 $0.95 / $4 13m04s
38 GLM 5.3 Flash Terminus 2 62.92% $0.03 $0.075 / $0.25 16m06s
39 Qwen 3.7 Max Terminus 2 61.05% $0.33 $2.5 / $7.5 15m35s
40 MiMo V2.5 Terminus 2 60.67% $0.02 $0.14 / $0.28 12m55s
41 Composer 2.5 Cursor CLI 58.43% N/A $0.5 / $2.5 N/A
42 Qwen 3.8 27B Terminus 2 58.43% $0.68 $0.5 / $3 14m53s
43 Gemini 4 Argon Terminus 2 57.68% N/A $4 / $20 14m02s
44 GPT 5.5 Codex 57.30% $0.65 N/A 3m22s
45 GPT 5.5 Factory 57.30% $0.75 N/A 6m29s
46 Claude Sonnet 4.6 Terminus 2 57.30% $0.57 $3 / $15 11m44s
47 MiMo V2.5 Pro Terminus 2 57.30% $0.04 $0.435 / $0.87 14m26s
48 GLM 5.1 Terminus 2 56.93% $0.23 $1 / $3.2 15m41s
49 Inkling Small Terminus 2 55.06% $0.13 $0.3 / $1.2 10m04s
50 Hy4 Preview Terminus 2 55.06% $0.14 $0.834 / $2.501 18m29s
51 GPT 5.4 Mini Terminus 2 54.68% $0.29 $0.75 / $4.5 11m13s
52 DeepSeek V4 Pro 0813 Terminus 2 54.68% $0.20 $1.32 / $3.96 14m22s
53 Gemini 3 Flash (12/25) Terminus 2 53.93% $0.15 $0.5 / $3 6m16s
54 Kimi K2.6 Terminus 2 53.56% $0.14 $0.95 / $4 15m32s
55 MiniMax-M3 Terminus 2 53.56% $0.21 $0.6 / $2.4 17m03s
56 Qwen 3.6 Plus Terminus 2 53.18% $0.20 $0.5 / $3 12m58s
57 Qwen 3.7 Plus Terminus 2 52.81% $0.10 $0.4 / $1.6 10m55s
58 Nemotron 3 Ultra Terminus 2 50.94% N/A N/A 9m44s
59 Gemini 3.5 Flash Lite Terminus 2 50.19% $0.34 $0.3 / $2.5 6m56s
60 Ling 3.0 Flash Terminus 2 50.19% $0.07 $0.075 / $0.22 9m54s
61 Ling 3.0 Flash Fin Terminus 2 50.19% $0.04 $0.06 / $0.18 13m54s
62 DeepSeek V4 Terminus 2 50.19% $0.24 $1.32 / $3.96 17m58s
63 MiniMax-M2.7 Terminus 2 48.69% $0.07 $0.3 / $1.2 13m33s
64 Inkling Terminus 2 47.57% $0.63 $1 / $4.05 12m52s
65 Grok 4.20 (Reasoning) Terminus 2 44.20% $0.41 $2 / $6 5m06s
66 Claude Haiku 4.5 (Thinking) Terminus 2 43.82% $0.36 $1 / $5 10m49s
67 Grok 4.3 Terminus 2 41.95% $0.98 $1.25 / $2.5 7m51s
68 Kimi K2.5 Terminus 2 41.95% $0.09 $0.6 / $3 17m31s
69 GPT 5.4 Nano Terminus 2 41.57% $0.07 $0.2 / $1.25 12m37s
70 Mistral Medium 3.5 Terminus 2 38.95% $1.19 $1.5 / $7.5 16m29s
71 Mercury 2.5 Terminus 2 34.46% $0.21 $0.2 / $0.75 11m29s
72 Gemini 3.1 Flash Lite Preview Terminus 2 34.08% $0.08 $0.25 / $1.5 4m38s
73 Laguna M.1 Terminus 2 34.08% N/A N/A 20m10s
74 Laguna XS.2 Terminus 2 25.84% N/A N/A 20m19s
75 Command A+ Terminus 2 17.60% N/A N/A N/A
76 Nemotron 3.5 Lightning Terminus 2 10.86% $0.01 $0.05 / $0.2 17m08s

Key Takeaways

  • Claude Opus 5.5 leads at 87.64%, a fraction ahead of GPT-6 Astra (87.27%), then GPT-5.6 Sol (85.77%), Claude Fable 5.1 (85.02%) and Claude Opus 5 (84.64%). The two leaders get there differently: Opus 5.5 takes the medium split at 90.30% but sits fifth on hard at 81.11%, while Astra’s edge is the hard split, where it reaches 87.78% against 84.44% for Sol and Opus 5 and 82.22% for Fable 5.1, with a sharp cliff to sixth place (74.44%); on medium Astra only ties Sol and Claude Sonnet 5.5 at 86.06%. While seventeen models reach 100% on easy, the median score on that split is 83.33%, so even the easy tasks still separate the field.
  • Note on the headline scores: Opus 5.5’s run used Opus 5 and Claude Opus 4.8 as fallbacks on 26 of 267 tasks; counting those as failures lowers it from 87.64% to 79.77%, behind Astra. Both Fable 5’s and Opus 5’s runs also used Opus 4.8 as a refusal fallback; counting Opus 5’s nine affected passes as failures lowers it from 84.64% to 81.27% (toggle the split icon beside the score to compare).

Background

Terminal-Bench 2.1 is an open-source benchmark that is designed to test a model’s ability to navigate and complete tasks in a sandboxed terminal environment. This version of the benchmark uses 89 tasks with unique categories ranging from model training to system administration. These tasks scale in difficulty from easy to hard.

We chose to include Terminal-Bench because a) it is increasingly common for it to be reported by model providers, b) it reflects the real-world terminal tasks expected of software engineers, and c) it is quite challenging, with no model scoring above 50% on the hard tasks upon its initial release. Furthermore, agentic systems like Claude Code, Codex, and Cursor now rely heavily on executing terminal commands correctly.

This benchmark was developed by the Terminal-Bench community as an open-source effort; we’d like to thank the community for their efforts in building this benchmark and for helping us integrate it into our evaluation suite. If you’re interested in learning more about Terminal-Bench or want to contribute to the project, visit tbench.ai

Below is an example task (you can find the full details for this task in the open-source task registry).

Please train a fasttext model on the yelp data in the data/ folder.

The final model size needs to be less than 150MB but get at least 0.62 accuracy on a private test set that comes from the same yelp review distribution.

The model should be saved as /app/model.bin


Methodology

All models were benchmarked using the Terminus 2 harness. Unless otherwise specified, we use identical configuration and methodology to Terminus 2. All results reported are pass@1.

On submission, we run the model against the provided pytests—a model must pass all pytests to get any credit for a task.

Unlike Terminus 1, Terminus 2 does not use structured outputs to enforce a response schema. Model queries returning an invalid or missing JSON are retried with a warning.


Comparison to Original Terminal-Bench

This benchmark is similar in structure to the original Terminal-Bench - it features 80 terminal-based tasks, many of which (like our example!) also appear in Terminal-Bench 2.1. The most substantial differences are in evaluation methododology:

  • We ran the original Terminal-Bench on an EC2 instance, using docker containers. Following the Laude implementation, we run Terminal-Bench 2.1 remotely using daytona.
  • We ran the original Terminal-Bench using a turn limit. Again, following the Laude implementation, we run Terminal-Bench 2.1 using a time limit instead.