Public

Effect-TS / effect

Benchmark updated: 9/15/202630 Tasks
Add models

Build production-ready applications in TypeScript

Languages

TypeScript99.9%JavaScript0.1%Shell<0.1%Nix<0.1%

Harness

1

Mini-SWE-agent
28 / 30

$6.14

16m28s

2

Mini-SWE-agent
27 / 30

$0.14

10m30s

Key Takeaways

  • Claude Opus 5 with Mini-SWE-agent scores 93.33% versus 90% for Deepseek V4 Flash 0731 with Mini-SWE-agent.
  • Deepseek V4 Flash 0731 with Mini-SWE-agent costs $0.14 per test versus $6.14 for Claude Opus 5 with Mini-SWE-agent.
  • With 30 tasks, these results are directional rather than a smoke test.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write
Claude Opus 5
7.7M
DeepSeek V4 Flash 0731
4.4M

Cost is the clearest tradeoff in this comparison. Claude Opus 5 leads at 93.33% for $6.14 per test. DeepSeek V4 Flash 0731 is the lower-cost option at 90.00% for $0.14 per test.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
Claude Opus 5
16m 28s
DeepSeek V4 Flash 0731
10m 30s

Latency separates several models with similarly strong scores. Claude Opus 5 leads at 93.33%, while DeepSeek V4 Flash 0731 is fastest at 10m 30s with 90.00% accuracy.