Public

gtrght / omnigent

Benchmark updated: 9/16/202630 Tasks
Add models

Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.

Languages

Python79.7%TypeScript17.6%JavaScript1.4%Kotlin0.3%Swift0.3%Rust0.3%Other0.5%

Harness

1

Mini-SWE-agent
30 / 30

$2.17

10m09s

Key Takeaways

  • This 30-task result is a complete benchmark run, not a directional or smoke-test result.
  • GPT-5.6 Sol with Mini-SWE-agent averages 14,360.10 output tokens and 13,669.73 reasoning tokens per task.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write
GPT-5.6 Sol
2.9M

Cost is the clearest tradeoff in this comparison. GPT-5.6 Sol leads at 100.00% for $2.17 per test. No other model in this comparison is cheaper.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
GPT-5.6 Sol
10m 9s

GPT-5.6 Sol is both the most accurate and fastest model in this comparison at 100.00% and 10m 9s.