Public
punt-labs / lux
Benchmark updated: 9/14/202630 Tasks
GUI - a whiteboard and dashboard interface for agents to share visual data with humans.
Languages
Python98.3%Shell0.8%TeX0.8%Makefile0.1%
Harness | Input / Output Cost | ||||||
|---|---|---|---|---|---|---|---|
1 | 22 / 30 | $1.19 | $1.25/$4.25 | 7m36s | |||
2 | 22 / 30 | $2.20 | N/A | 20m07s | |||
3 | 18 / 30 | $0.82 | $1.4/$4.4 | 11m23s |
Key Takeaways
- This 30-task result is directional: Claude Opus 5 (High Effort) and Muse Spark 1.2 each score 73.33%, while GLM 5.2 (Fireworks) scores 60%.
- Muse Spark 1.2 with Mini-SWE-agent has lower cost per test and latency than Claude Opus 5 (High Effort): $1.19 and 456 seconds versus $2.20 and 1207 seconds.
- GLM 5.2 (Fireworks) with Mini-SWE-agent is the lowest-cost run at $0.82 per test, but resolves four fewer tasks than the two leading runs.
Cost Analysis
Cost / Test vs. Accuracy
ACCURACYCOST
Average Token Use / Test
Token Usage
InputOutputReasoningCache readCache write
Muse Spark 1.2
GLM 5.2
anthropic/claude-opus-5-high
Cost is the clearest tradeoff in this comparison. anthropic/claude-opus-5-high leads at 73.33% for $2.20 per test. Muse Spark 1.2 is the lower-cost option at 73.33% for $1.19 per test.
Latency Analysis
Latency vs. Accuracy
ACCURACYLATENCY
Average Response Time / Test
Response Time
anthropic/claude-opus-5-high
GLM 5.2
Muse Spark 1.2
Latency separates several models with similarly strong scores. anthropic/claude-opus-5-high leads at 73.33%, while Muse Spark 1.2 is fastest at 7m 36s with 73.33% accuracy.