Release Date: Jul 15, 2026

Developer Thinking Machinesย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 256k
Max Output Tokens 256k
Token Costs (in/out) $1.00/4.05
Weights Open
Input Modalities

Accuracy

34.10 % ยฑ 1.13

Cost / Test (Vals Index)

$ 1.354

Latency

19 min 37 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.13
33/46

0.0%

ยฑ3.40
37/49

0.0%

ยฑ6.38
15/22

0.0%

ยฑ3.27
34/46

0.0%

ยฑ0.93
26/49

0.0%

ยฑ3.13
23/49

0.0%

ยฑ2.23
41/85

0.0%

ยฑ1.84
15/84

0.0%

ยฑ0.94
48/95

0.0%

ยฑ0.00
23/23

0.0%

ยฑ3.22
52/76

0.0%

ยฑ1.28
20/30

0.0%

ยฑ0.84
18/140

0.0%

ยฑ3.66
58/84

0.0%

ยฑ1.78
36/133

0.0%

ยฑ1.00
30/138

0.0%

ยฑ0.50
53/137

0.0%

ยฑ0.34
40/133

0.0%

ยฑ0.99
47/89

0.0%

ยฑ0.00
36/39

0.0%

ยฑ3.86
28/28

0.0%

ยฑ1.87
29/83

0.0%

ยฑ0.38
45/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Thinkingmachines
Temperature: 1
Top P: 1
Top K: Default
Max Output Tokens: 256,000
Reasoning Effort: 0.99

Updates

Jul 15, 2026

We evaluated Thinking Machinesโ€™ first open-weights model, Inkling, on the Vals Index.

  • Inkling scores 34.10% on the Vals Index, placing #30 of 43 models overall and #9 among open-weight models.

  • Its strongest component results include 75.49% on the SWE-bench Verified Vals Index subset and 69.23% on the CorpFin v2 Vals Index subset. It also scores 45.97% on the Finance Agent v2 Index subset.

  • Inkling scores 13.28% on the Vibe Code Bench Vals Index subset and 47.57% across three full trials of Terminal-Bench 2.1.

We used Thinking Machinesโ€™ recommended evaluation settings: temperature 1, top-p 1, up to 256k output tokens, separate reasoning enabled, and reasoning effort set to "0.99". Inkling supports a 256k context window, image inputs, and tool calling.

Congrats to the Thinking Machines team on the release!