Public
gtrght / omnigent
Benchmark updated: 9/16/202630 Tasks
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
Languages
Python79.7%TypeScript17.6%JavaScript1.4%Kotlin0.3%Swift0.3%Rust0.3%Other0.5%
Harness | Input / Output Cost | ||||||
|---|---|---|---|---|---|---|---|
1 | 30 / 30 | $2.17 | $4/$20 | 10m09s |
Key Takeaways
- This 30-task result is a complete benchmark run, not a directional or smoke-test result.
- GPT-5.6 Sol with Mini-SWE-agent averages 14,360.10 output tokens and 13,669.73 reasoning tokens per task.
Cost Analysis
Cost / Test vs. Accuracy
ACCURACYCOST
Average Token Use / Test
Token Usage
InputOutputReasoningCache readCache write
GPT-5.6 Sol
Cost is the clearest tradeoff in this comparison. GPT-5.6 Sol leads at 100.00% for $2.17 per test. No other model in this comparison is cheaper.
Latency Analysis
Latency vs. Accuracy
ACCURACYLATENCY
Average Response Time / Test
Response Time
GPT-5.6 Sol
GPT-5.6 Sol is both the most accurate and fastest model in this comparison at 100.00% and 10m 9s.