Public
vastsa / PI-Desktop
Benchmark updated: 9/10/202630 Tasks
Local-first AI coding agent desktop: Electron + Rust host core + pi Agent Harness + user-installable plugins
Languages
TypeScript57.3%Rust18.2%JavaScript17.6%CSS6%HTML0.8%Python0.1%Other<0.1%
Harness | Input / Output Cost | ||||||
|---|---|---|---|---|---|---|---|
1 | 12 / 30 | $0.87 | $1.25/$4.25 | 5m52s | |||
2 | 11 / 30 | $0.07 | $0.44/$1.32 | 6m01s | |||
3 | 10 / 30 | $0.19 | $0.4/$1.6 | 2m38s |
Key Takeaways
- This 30-task evaluation is directional: Muse Spark 1.2 with Mini-SWE-agent scores 40%, ahead of Deepseek V4 Flash 0731 at 36.67% and Qwen3.7 Plus at 33.33%.
- Qwen3.7 Plus with Mini-SWE-agent is fastest at 158.42 seconds per test, while Deepseek V4 Flash 0731 is least expensive at $0.07 per test.
Cost Analysis
Cost / Test vs. Accuracy
ACCURACYCOST
Average Token Use / Test
Token Usage
InputOutputReasoningCache readCache write
Muse Spark 1.2
DeepSeek V4 Flash 0731
Qwen 3.7 Plus
Cost is the clearest tradeoff in this comparison. Muse Spark 1.2 leads at 40.00% for $0.87 per test. DeepSeek V4 Flash 0731 is the lower-cost option at 36.67% for $0.07 per test.
Latency Analysis
Latency vs. Accuracy
ACCURACYLATENCY
Average Response Time / Test
Response Time
DeepSeek V4 Flash 0731
Muse Spark 1.2
Qwen 3.7 Plus
Latency separates several models with similarly strong scores. Muse Spark 1.2 leads at 40.00%, while Qwen 3.7 Plus is fastest at 2m 38s with 33.33% accuracy.