Public
Effect-TS / effect
Benchmark updated: 9/15/202630 Tasks
Build production-ready applications in TypeScript
Languages
TypeScript99.9%JavaScript0.1%Shell<0.1%Nix<0.1%
Harness | Input / Output Cost | ||||||
|---|---|---|---|---|---|---|---|
1 | 28 / 30 | $6.14 | $5/$25 | 16m28s | |||
2 | 27 / 30 | $0.14 | $0.44/$1.32 | 10m30s |
Key Takeaways
- Claude Opus 5 with Mini-SWE-agent scores 93.33% versus 90% for Deepseek V4 Flash 0731 with Mini-SWE-agent.
- Deepseek V4 Flash 0731 with Mini-SWE-agent costs $0.14 per test versus $6.14 for Claude Opus 5 with Mini-SWE-agent.
- With 30 tasks, these results are directional rather than a smoke test.
Cost Analysis
Cost / Test vs. Accuracy
ACCURACYCOST
Average Token Use / Test
Token Usage
InputOutputReasoningCache readCache write
Claude Opus 5
DeepSeek V4 Flash 0731
Cost is the clearest tradeoff in this comparison. Claude Opus 5 leads at 93.33% for $6.14 per test. DeepSeek V4 Flash 0731 is the lower-cost option at 90.00% for $0.14 per test.
Latency Analysis
Latency vs. Accuracy
ACCURACYLATENCY
Average Response Time / Test
Response Time
Claude Opus 5
DeepSeek V4 Flash 0731
Latency separates several models with similarly strong scores. Claude Opus 5 leads at 93.33%, while DeepSeek V4 Flash 0731 is fastest at 10m 30s with 90.00% accuracy.