Public

TuringWorks / orbit-rs

Benchmark updated: 9/15/202630 Tasks
Add models

Orbit-RS: A performant Multi-Protocol Data Platform

Languages

Rust88.2%ANTLR2.9%Python2.1%PLpgSQL1.5%Shell1.3%JavaScript1.2%Other2.8%

Harness

1

Mini-SWE-agent
25 / 30

$3.02

1h04m

2

Mini-SWE-agent
17 / 30

$0.35

41m11s

Key Takeaways

  • Claude Sonnet 5 with Mini-SWE-agent scores 83.33%, compared with 56.67% for Claude Haiku 4.5 (Nonthinking) with Mini-SWE-agent.
  • Claude Sonnet 5 with Mini-SWE-agent averages $3.02 and 3856.06 seconds per test; Claude Haiku 4.5 (Nonthinking) with Mini-SWE-agent averages $0.35 and 2470.92 seconds.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write

No token usage data available.

Cost is the clearest tradeoff in this comparison. Claude Sonnet 5 leads at 83.33% for $3.02 per test. Claude Haiku 4.5 (Nonthinking) is the lower-cost option at 56.67% for $0.35 per test.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
Claude Sonnet 5
64m 16s
Claude Haiku 4.5 (Nonthinking)
41m 11s

Latency separates several models with similarly strong scores. Claude Sonnet 5 leads at 83.33%, while Claude Haiku 4.5 (Nonthinking) is fastest at 41m 11s with 56.67% accuracy.