Public

vastsa / PI-Desktop

Benchmark updated: 9/10/202630 Tasks
Add models

Local-first AI coding agent desktop: Electron + Rust host core + pi Agent Harness + user-installable plugins

Languages

TypeScript57.3%Rust18.2%JavaScript17.6%CSS6%HTML0.8%Python0.1%Other<0.1%

Harness

1

Mini-SWE-agent
12 / 30

$0.87

5m52s

2

Mini-SWE-agent
11 / 30

$0.07

6m01s

3

Mini-SWE-agent
10 / 30

$0.19

2m38s

Key Takeaways

  • This 30-task evaluation is directional: Muse Spark 1.2 with Mini-SWE-agent scores 40%, ahead of Deepseek V4 Flash 0731 at 36.67% and Qwen3.7 Plus at 33.33%.
  • Qwen3.7 Plus with Mini-SWE-agent is fastest at 158.42 seconds per test, while Deepseek V4 Flash 0731 is least expensive at $0.07 per test.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write
Muse Spark 1.2
4.3M
DeepSeek V4 Flash 0731
2.1M
Qwen 3.7 Plus
1.7M

Cost is the clearest tradeoff in this comparison. Muse Spark 1.2 leads at 40.00% for $0.87 per test. DeepSeek V4 Flash 0731 is the lower-cost option at 36.67% for $0.07 per test.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
DeepSeek V4 Flash 0731
6m 1s
Muse Spark 1.2
5m 52s
Qwen 3.7 Plus
2m 38s

Latency separates several models with similarly strong scores. Muse Spark 1.2 leads at 40.00%, while Qwen 3.7 Plus is fastest at 2m 38s with 33.33% accuracy.