Release Date: Jul 15, 2026

Developer Thinking Machinesย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 256k
Max Output Tokens 256k
Token Costs (in/out) $1.00/4.05
Weights Open
Input Modalities

Accuracy

34.10 % ยฑ 1.13

Cost / Test (Vals Index)

$ 1.354

Latency

19 min 37 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.13
34/48

0.0%

ยฑ3.40
38/51

0.0%

ยฑ6.38
15/22

0.0%

ยฑ3.27
35/48

0.0%

ยฑ0.93
27/51

0.0%

ยฑ3.13
24/51

0.0%

ยฑ2.23
41/86

0.0%

ยฑ1.84
15/87

0.0%

ยฑ0.94
49/96

0.0%

ยฑ0.00
24/24

0.0%

ยฑ3.22
53/77

0.0%

ยฑ1.28
20/30

0.0%

ยฑ0.84
18/141

0.0%

ยฑ3.66
59/86

0.0%

ยฑ1.78
38/135

0.0%

ยฑ1.00
30/140

0.0%

ยฑ0.50
54/139

0.0%

ยฑ0.34
41/135

0.0%

ยฑ0.99
48/90

0.0%

ยฑ0.00
39/42

0.0%

ยฑ3.86
30/30

0.0%

ยฑ1.87
32/86

0.0%

ยฑ0.38
47/57
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Thinkingmachines
Temperature: 1
Top P: 1
Top K: Default
Max Output Tokens: 256,000
Reasoning Effort: 0.99

Updates

Jul 15, 2026

We evaluated Thinking Machinesโ€™ first open-weights model, Inkling, on the Vals Index.

  • Inkling scores 34.10% on the Vals Index, placing #30 of 43 models overall and #9 among open-weight models.

  • Its strongest component results include 75.49% on the SWE-bench Verified Vals Index subset and 69.23% on the CorpFin v2 Vals Index subset. It also scores 45.97% on the Finance Agent v2 Index subset.

  • Inkling scores 13.28% on the Vibe Code Bench Vals Index subset and 47.57% across three full trials of Terminal-Bench 2.1.

We used Thinking Machinesโ€™ recommended evaluation settings: temperature 1, top-p 1, up to 256k output tokens, separate reasoning enabled, and reasoning effort set to "0.99". Inkling supports a 256k context window, image inputs, and tool calling.

Congrats to the Thinking Machines team on the release!