Vals-Smith
/ apache / spark
Key Takeaways
- GPT-5.6 Sol leads with 49 of 58 tasks resolved and uniquely solves two tasks.
- Claude Fable 5 and Claude Opus 4.8 follow at 48, while Grok 4.5 is the strongest value alternative at 47 tasks for $0.71 per test.
- GPT-5.6 Sol and Claude Fable 5 resolve 54 of 58 tasks together, capturing every success found across the model set.
Cost / Test vs. Accuracy
Grok 4.5 resolves 47 tasks for $0.71 per test, while GPT-5.6 Sol resolves 49 for $2.46.
Latency vs. Accuracy
Grok 4.5 reaches 47 of 58 tasks in 227 seconds, compared with 49 tasks in 588 seconds for GPT-5.6 Sol.
Average token use / test
Input
Output
Reasoning
Cache read
Cache write
Claude Sonnet 5
Claude Opus 4.7
Claude Opus 4.8
GLM 5.2 (Fireworks)
Claude Fable 5
Kimi K3
Gemini 3.5 Flash
GPT-5.6 Luna
Claude Haiku 4.5 (Thinking)
GPT-5.6 Sol
GPT 5.5
Muse Spark 1.1
Grok 4.5
Gemini 3.1 Pro Preview (02/26)
GPT-5.6 Terra
Tasks with failures
PassedFailed
| Model | 12226c4 | 1cd649d | 1db5e68 | 217fb4e | 2274217 | 233b158 | 308de00 | 384d5e8 | 475a771 | 484342a | 50932d2 | 64876d2 | 653eb49 | 6b24038 | 6db4ab9 | 715a437 | 7dc70c1 | 88359ec | 9d760a9 | a3bb3f7 | a8265e3 | af739dd | b22a865 | babd0ee | c07d9be | c1fb0aa | c76b9e2 | c9c45a6 | cd3b583 | d8b8f39 | ea9d6a7 | ec7173c | ed150f9 | f2f6bde |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol | ||||||||||||||||||||||||||||||||||
| Claude Opus 4.8 | ||||||||||||||||||||||||||||||||||
| Claude Fable 5 | ||||||||||||||||||||||||||||||||||
| Grok 4.5 | ||||||||||||||||||||||||||||||||||
| GLM 5.2 | ||||||||||||||||||||||||||||||||||
| Kimi K3 | ||||||||||||||||||||||||||||||||||
| Claude Sonnet 5 | ||||||||||||||||||||||||||||||||||
| Claude Opus 4.7 | ||||||||||||||||||||||||||||||||||
| GPT-5.6 Luna | ||||||||||||||||||||||||||||||||||
| Muse Spark 1.1 | ||||||||||||||||||||||||||||||||||
| GPT-5.6 Terra | ||||||||||||||||||||||||||||||||||
| Gemini 3.5 Flash | ||||||||||||||||||||||||||||||||||
| GPT 5.5 | ||||||||||||||||||||||||||||||||||
| Claude Haiku 4.5 (Thinking) | ||||||||||||||||||||||||||||||||||
| Gemini 3.1 Pro Preview (02/26) |
Head-to-head
VS
Accuracy
84.48%Δ 1.72%
82.76%
Cost / test
$2.46Δ $2.09
$4.55
Latency
588sΔ 450s
1038s
Cost distribution
Latency distribution
Input token distribution
Output token distribution
Task outcomes
58 tasks
Both
A only
B only
Neither
Both: 44 tasks
A only: 5 tasks
B only: 4 tasks
Neither: 5 tasks