oven-sh/bun
Languages
Incredibly fast JavaScript runtime, bundler, test runner, and package manager – all in one.
Harness | Input / Output Cost | ||||||
|---|---|---|---|---|---|---|---|
1 | 26 / 48 | $2.26 | $3/$15 | 16m45s | |||
2 | 25 / 48 | $4.11 | $5/$25 | 15m57s | |||
3 | 25 / 48 | $7.33 | $10/$50 | 17m16s | |||
4 | 24 / 48 | $2.29 | $5/$30 | 8m34s | |||
5 | 24 / 48 | $4.00 | $5/$25 | 13m00s | |||
6 | 23 / 48 | $0.81 | $2/$12 | 4m12s | |||
7 | 23 / 48 | $0.11 | $0.2/$1.2 | 4m59s | |||
8 | 23 / 48 | $1.80 | $5/$30 | 6m01s | |||
9 | 23 / 48 | $0.59 | $2/$6 | 7m42s | |||
10 | 22 / 48 | $4.65 | $3/$15 | 21m07s | |||
11 | 21 / 48 | $0.38 | $2/$12 | 2m37s | |||
12 | 21 / 48 | $1.24 | $1.5/$9 | 5m27s | |||
13 | 21 / 48 | $0.74 | $1.4/$4.4 | 15m37s | |||
14 | 16 / 48 | $0.30 | $1/$5 | 4m13s |
Key Takeaways
- GPT-5.6 Luna with Mini-SWE-agent resolves 23 of 48 tasks for $0.11 per test, the lowest cost among the supplied runs.
- GPT-5.6 Sol with Mini-SWE-agent resolves 24 of 48 tasks in 514 seconds, matching Claude Opus 4.7 on tasks resolved but finishing faster.
- Claude Haiku 4.5 (Nonthinking) with Mini-SWE-agent resolves 16 of 48 tasks, the lowest score in this set.
Model Comparison
Accuracy
54.17%
Kimi K3
52.08%
Claude Opus 4.8
Task outcomes
48 tasks
Cost / test
$2.26
Kimi K3
$4.11
Claude Opus 4.8
Cost distribution
Latency
16m 45s
Kimi K3
15m 57s
Claude Opus 4.8
Latency distribution
Cost Analysis
Average Token Use / Test
Cost is the clearest tradeoff in this comparison. Kimi K3 leads at 54.17% for $2.26 per test. No other model in this comparison is cheaper.
Latency Analysis
Average Response Time / Test
Latency separates several models with similarly strong scores. Kimi K3 leads at 54.17%, while GPT-5.6 Terra is fastest at 2m 37s with 43.75% accuracy.
Tasks with failures
| Models | |||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Kimi K3 | |||||||||||||||||||||||||||||||||||
| Claude Opus 4.8 | |||||||||||||||||||||||||||||||||||
| Claude Fable 5 | |||||||||||||||||||||||||||||||||||
| GPT-5.6 Sol | |||||||||||||||||||||||||||||||||||
| Claude Opus 4.7 | |||||||||||||||||||||||||||||||||||
| GPT-5.6 Luna | |||||||||||||||||||||||||||||||||||
| Grok 4.5 | |||||||||||||||||||||||||||||||||||
| Gemini 3.1 Pro Preview (02/26) | |||||||||||||||||||||||||||||||||||
| GPT 5.5 | |||||||||||||||||||||||||||||||||||
| Claude Sonnet 5 | |||||||||||||||||||||||||||||||||||
| GPT-5.6 Terra | |||||||||||||||||||||||||||||||||||
| GLM 5.2 | |||||||||||||||||||||||||||||||||||
| Gemini 3.5 Flash | |||||||||||||||||||||||||||||||||||
| Claude Haiku 4.5 (Nonthinking) |
Task detail
4bbe075Issue statement
Bun's mimalloc dependency definition still applies a local workaround for an out-of-bounds strnlen read even though the pinned mimalloc revision already includes the upstream correction. Remove the obsolete patch from the mimalloc dependency configuration so source acquisition uses the pinned upstream archive without trying to apply that redundant patch.
View Hidden Tests
diff --git a/test/internal/mimalloc-dependency.test.ts b/test/internal/mimalloc-dependency.test.tsnew file mode 100644index 0000000000..481ba6fe13--- /dev/null+++ b/test/internal/mimalloc-dependency.test.ts@@ -0,0 +1,10 @@+import { expect, test } from "bun:test";++import { mimalloc } from "../../scripts/build/deps/mimalloc.ts";++test("pinned mimalloc source needs no local patches", () => {+ const source = mimalloc.source({} as never);++ expect(source.kind).toBe("github-archive");+ expect(mimalloc.patches).toBeUndefined();+});