Research

Do Agent Teams Pay Off? A Case Study on Vibe Code Bench

We ran GPT 6 Sol and Claude Opus 5.5 on Vibe Code Bench as single agents and as teams, at medium and max reasoning effort. Teams cost 1.8 to 5.1 times as much, and only for Sol at medium effort was the difference in score statistically significant. Whether a team is worth it depends on the model, task, and budget: a team can boost performance, but higher reasoning effort is an alternative worth weighing.

Katrina Drozdov 10/09/2026
Do Agent Teams Pay Off? A Case Study on Vibe Code Bench

Both OpenAI and Anthropic now ship native multi-agent support: a lead agent can spawn subagents, hand each of them part of a task, and assemble what they return. We wanted to measure what that buys, what it costs, and which kinds of work it helps, on tasks that resemble real software projects.

We used Vibe Code Bench, our benchmark for building complete web applications from a product spec, as a test bed. Each app is deployed with Docker Compose and graded by browser agents that click through its workflows. The work splits naturally into a database, an API, an interface, and deployment, and bugs at the seams between those pieces show up directly in the score. We ran GPT 6 Sol and Claude Opus 5.5 on 50 apps, each as a single agent and as a team, at medium and max reasoning effort.

In a team, the lead agent can hand parts of the build to subagents, up to five at once, and subagents cannot spawn their own. Both leads received the same short instruction: split the work into independent parts, delegate each one, then integrate and verify the result; the task prompt, sandbox, and grader were otherwise identical. Score is the share of an app’s UI tests that pass, averaged over the 50 apps. Each of the eight setups ran once.

Key findings

  • Teams cost 1.8 to 5.1 times as much as single agents, and only one difference in score was statistically significant. Sol’s medium-effort team gained 7.3 points over its single agent (p = 0.005). The other three differences, from −0.3 to +3.4 points, were not statistically significant.
  • For Sol, higher effort raised the score more than a team did; for Opus, neither higher effort nor a team made a significant difference. Higher effort came at a higher price, though: Sol’s max-effort single agent cost 1.6 times as much as its medium-effort team ($4.82 vs. $3.07 per app) and took three times as long.
  • Teams finished faster for Sol but not for Opus, and the two leads orchestrated differently. Sol’s lead split the app across parallel subagents within the first minutes and tested the result itself; at max effort, its runtime fell from 38 to 28 minutes. Opus’s lead usually first wrote a shared contract for how the app’s parts would fit together, then delegated in sequential waves of builders and testers, and its teams took up to 2.3 times as long as its single agents.

These observations come from one benchmark of full-stack web apps and may not carry over to other kinds of work.

Cost, runtime, and score

We compared each pair on the same 50 apps with a paired t-test. Only Sol’s medium-effort gain, 7.3 points, is statistically significant (p = 0.005; 95% interval +2.3 to +12.4 points). The other three are not (p = 0.16 to 0.81): +1.4 points for Sol at max effort, −0.3 for Opus at medium, and +3.4 for Opus at max, with intervals of −1.7 to +4.5, −3.2 to +2.5, and −1.4 to +8.3 points.

Cost per app vs. score on Vibe Code Bench
GPT 6 SolClaude Opus 5.5Single agentTeam

Each arrow points from a model's single agent to its team at the same reasoning effort.

Score, cost per app, and runtime by model, effort, and setup
ModelSetupScoreMean cost per appMedian runtime
GPT 6 Solmedium effort, single agent77.6%$1.2212.4 min
GPT 6 Solmedium effort, team84.9%$3.0711.9 min
GPT 6 Solmax effort, single agent89%$4.8238.4 min
GPT 6 Solmax effort, team90.4%$8.5427.6 min
Claude Opus 5.5medium effort, single agent91.5%$4.0820.3 min
Claude Opus 5.5medium effort, team91.2%$9.1026.6 min
Claude Opus 5.5max effort, single agent89.8%$23.7774.7 min
Claude Opus 5.5max effort, team93.2%$122174.1 min

At medium and max effort respectively, teams cost 2.5 and 1.8 times as much as the single agent for Sol, and 2.2 and 5.1 times as much for Opus. Most of the extra spend is cached input: each subagent has its own context holding the spec and its working state, and that context is re-sent on every model call. At max effort, Opus’s subagents read a median of 224 million cached tokens per app, against 17 million for the lead and 55 million for the single agent. One Opus max-effort team ran out of time while still verifying; its cost of about $189 is estimated from its token counts and included in the mean.

Runtime vs. score on Vibe Code Bench
GPT 6 SolClaude Opus 5.5Single agentTeam

Each arrow points from a model's single agent to its team at the same reasoning effort.

Score, cost per app, and runtime by model, effort, and setup
ModelSetupScoreMean cost per appMedian runtime
GPT 6 Solmedium effort, single agent77.6%$1.2212.4 min
GPT 6 Solmedium effort, team84.9%$3.0711.9 min
GPT 6 Solmax effort, single agent89%$4.8238.4 min
GPT 6 Solmax effort, team90.4%$8.5427.6 min
Claude Opus 5.5medium effort, single agent91.5%$4.0820.3 min
Claude Opus 5.5medium effort, team91.2%$9.1026.6 min
Claude Opus 5.5max effort, single agent89.8%$23.7774.7 min
Claude Opus 5.5max effort, team93.2%$122174.1 min

Parallelism saved time for Sol at both efforts, slightly at medium: its team finished in 11.9 minutes against 12.4 for its single agent at medium effort, and in 28 against 38 at max.

Opus’s teams took longer, up to 2.3 times the single agent’s runtime at max effort. They worked in sequential waves, described below, with each wave waiting for its slowest member, and at max effort the median subagent ran for 46 minutes, against 7 at medium.

The extra time was subagent work, not idle time, and longer runs did not score higher. Across apps, the Opus max team’s runtime had no correlation with its gain over the single agent. Four of its runs took more than four hours, and one hit the 5.25-hour limit.

Higher effort or a team

A team and higher effort both raise cost, and their effects on score differed. For Sol, max effort raised the single agent’s score by 11.4 points (p < 0.001), more than the 7.3 points a team added at medium effort. The max-effort single agent also scored 4.1 points above the medium-effort team (89.0% vs. 84.9%), a difference that was not statistically significant (p = 0.08), at 1.6 times the cost and three times the runtime.

For Opus, max effort alone did not raise the score: the single agent scored 89.8% for $23.77, against 91.5% for $4.08 at medium (p = 0.48). Max effort combined with a team produced the highest score in the study, 93.2%, at thirty times the cost of the medium-effort single agent; its 1.6-point lead over that agent was not statistically significant (p = 0.36). Across models, Sol’s max-effort team (90.4%, $8.54) came within about a point of Opus’s medium-effort single agent, at twice the price.

Neither option was a clear default: higher effort gave Sol the larger gain at a higher price, while Opus’s cheapest setup scored about as well as its most expensive one on this set of tasks.

How the leads worked

The two leads received the same instruction and divided the work differently.

Sol’s lead split the app and did the testing itself

At medium effort, Sol’s lead split the work along architectural lines (database and seed data, API, interface, and deployment) and created all of its subagents within about a minute. At max effort it more often split the interface into separate feature areas (13 of 50 opening plans, against 7 at medium) and kept adding subagents later in the run. It also spawned slightly more subagents (5.1 vs. 4.2) and sent them almost twice as many follow-up requests (31 vs. 18 per app).

Sol’s logs record the lead’s commands but not its file edits, so we can’t tell how much code it wrote; its recorded actions were mostly coordinating, building, and testing. It drove a headless browser through the app in 46 of 50 runs at both efforts, running about twice as many browser tests as the single agent, mostly at the app’s public URL rather than at localhost inside the sandbox. At medium effort, the single agent browser-tested only at localhost in 18 of 50 runs, against 4 for the team lead. In its notes, the lead reported catching bugs that only appear when separately built pieces meet, such as an interface sending fields that didn’t match the database schema, or a total computed one way in the interface and another in the database. Each fix went back to the subagent that owned the file.

Opus’s lead wrote a contract first and delegated in waves

Before spawning any agents, Opus’s lead wrote a shared document, usually CONTRACT.md, describing the schema, API routes, and file ownership (42 of 50 runs at medium, 40 at max; the single agents almost never did). It spawned subagents in waves. At medium effort, it usually spawned builders and then, in most runs, a subagent dedicated to end-to-end testing. At max effort, it added a wave: first subagents that built the foundation, then subagents that built features on top of it, then subagents that tested and reviewed the app. It wrote the READMEs itself in most runs and edited application code in 10 of 50 runs at medium effort and 26 at max.

At max effort, the Opus team was larger: 6.8 subagents per app instead of 4.2, a testing subagent in all 50 runs, and about 1,140 subagent tool calls per app, against about 270 for the max-effort single agent. Its subagents ran 36 browser tests per app, six times as many as the max-effort single agent.

Where the time went

Teams and single agents reached their first deploy (the first docker compose up) at similar times. At medium effort, both models’ teams were within a minute of their single agents. At max effort, Sol’s team deployed about five minutes earlier and Opus’s about six minutes later. Opus’s max-effort team then spent two thirds of its run after the first deploy, against under a third for its single agent. Over that stretch, its lead and subagents put a median of 68 agent-minutes (time summed across the lead and its subagents) into testing, against 10 for the single agent.

Time before and after the first deploy
Building, until the first deployAfter the first deploy, GPT 6 SolAfter the first deploy, Claude Opus 5.5

Medians over 50 apps. Bars end at the median runtime and split at the median time of the first docker compose up.

Median minutes to first deploy and to finish
Model and effortSetupFirst deployFinishedAgent-minutes of testing after the first deploy
GPT 6 Sol, mediumSingle agent5.1 min12.4 min3.1 min
GPT 6 Sol, mediumTeam5.7 min11.9 min3.5 min
GPT 6 Sol, maxSingle agent19.9 min38.4 min8 min
GPT 6 Sol, maxTeam15.3 min27.6 min7 min
Claude Opus 5.5, mediumSingle agent12.2 min20.3 min3.6 min
Claude Opus 5.5, mediumTeam12.5 min26.6 min6 min
Claude Opus 5.5, maxSingle agent51 min74.7 min9.6 min
Claude Opus 5.5, maxTeam56.8 min174.1 min67.8 min

The timeline below shows one app through all eight setups.

What each agent was doing, minute by minuteAll eight setups build the same app: a web app with user accounts, record keeping, a usage dashboard, and email notifications.
Writing codeBuild and deployTestingCoordinatingReading and reasoningWaiting on subagentsSubagent assigned (activity not logged)Possibly still workingMessage from the lead

Sol's logs record the lead's commands but not edits made with its file-editing tool, so time spent on those edits counts toward the next logged action, usually a build. A Sol subagent's bar is solid from its creation to the lead's last message to it, then dashed until the lead's last wait for subagents ended. Activity labels are heuristic classifications of tool calls.

Minutes per activity for each agent, Sol, medium, a web app with user accounts, record keeping, a usage dashboard, and email notifications
RunAgentWriting codeBuild and deployTestingCoordinatingReading and reasoningWaiting on subagents
Single agent (score 50%, 10 min, first deploy at 4.0 min)Agent0.0 min4.5 min4.9 min0.0 min0.8 min0.0 min
Team (score 62%, 14 min, first deploy at 6.8 min)Lead0.0 min2.5 min3.1 min1.2 min1.6 min5.2 min
Team (score 62%, 14 min, first deploy at 6.8 min)Subagent 1Activity not logged. Created at 0.3 min; 2 messages from the lead, the last at 2.1 min; possibly working until 9.7 min, when the lead's last wait for subagents ended.
Team (score 62%, 14 min, first deploy at 6.8 min)Subagent 2Activity not logged. Created at 0.3 min; 6 messages from the lead, the last at 8.7 min; possibly working until 9.7 min, when the lead's last wait for subagents ended.
Team (score 62%, 14 min, first deploy at 6.8 min)Subagent 3Activity not logged. Created at 0.4 min; 6 messages from the lead, the last at 8.7 min; possibly working until 9.7 min, when the lead's last wait for subagents ended.
Team (score 62%, 14 min, first deploy at 6.8 min)Subagent 4Activity not logged. Created at 0.5 min; 3 messages from the lead, the last at 2.8 min; possibly working until 9.7 min, when the lead's last wait for subagents ended.

Sol’s team leads delegated within the first two minutes, spent about a third of the run waiting on subagents, and tested the app once it was deployed.

Opus’s max-effort team shows the waves. The lead spent its first quarter hour on the contract. Three builders worked until about an hour in, two feature agents until almost two hours, and a testing agent until the end, while the lead waited. Its single agent finished in about an hour.

On this app, Sol’s teams scored higher than its single agents at both efforts (62% vs. 50% and 75% vs. 62%), and Opus’s did not (75% vs. 88% and 75% vs. 75%).

Where team gains came from

To check whether teams helped more on harder apps, we grouped apps using a run outside each comparison. For the medium-effort comparison, an app counts as harder if the same model’s max-effort single agent scored below 90% on it; for the max-effort comparison, we use the medium-effort single agent. This keeps chance results in the compared runs from deciding an app’s group.

Sol’s medium-effort team gained on about half the apps, most on harder ones. It scored higher than its single agent on 24 apps, lower on 11, and the same on 15. On harder apps it added about 15 points; on easier ones, about 1. At max effort, Sol’s smaller team gain was spread evenly across harder and easier apps.

Where Sol’s medium-effort team gained 20 points or more, it built and checked what the single agent skipped. On these nine apps, the single agents reported that testing had passed even when some features were never built and whole workflows were never tried. The team covered more of the spec and checked its work more carefully: it built those features, and its lead reviewed the separately built parts against each other, catching mismatches such as a form that didn’t fit the database schema or a permission rule that blocked a step.

Opus’s gains and losses were closer to balanced and did not favor harder apps. Its team scored higher on 14 apps and lower on 15 at medium effort, and higher on 13 and lower on 8 at max. At medium effort, it lost about four points on harder apps and gained about two on easier ones; at max effort, it added about three to four points on both.

Other properties we checked had no consistent effect. Apps with more distinct user roles gained 15 points from Sol’s medium-effort team, but no other configuration showed this. Spec length, the number of UI tests, and single-user versus multi-user apps showed no consistent effect either.

Future work

On this benchmark, the effect of a team depended on the model, the reasoning effort, and the app. The open question is which tasks teams handle reliably better than single agents. A team should gain the most when the work splits into parts that can proceed side by side, and the lead can combine them without spending the time it saved on coordination. Tasks with many independent parts and long stretches of separable work, such as large codebases, broad research questions, or analyses spanning many documents, are where we would look next.

Methodology and limitations

GPT 6 Sol ran through the OpenAI Agents API, with the API’s multi_agent mode off for single agents and on for teams, and with commands executed in a Vals sandbox. Claude Opus 5.5 ran through the Claude Code CLI in headless mode, with the subagent tools removed for single agents and available for teams. Opus’s team runs also disabled Claude Code’s ten-minute limit on background subagents. Reasoning effort, medium or max, was set through OpenAI’s reasoning.effort parameter and Claude Code’s CLAUDE_CODE_EFFORT_LEVEL. Apps were scored by the standard Vibe Code Bench UI test suite.

Behavioral counts come from each lead’s event stream, and from subagents’ streams for Opus. OpenAI’s Agents API emits two subagent.created events per subagent, so we counted unique subagent IDs for Sol. Browser tests count commands that launch or navigate a headless browser. Runtime is the agent’s wall-clock time in the sandbox. The harnesses expose different amounts of detail. Claude Code logs every subagent’s tool calls. OpenAI’s Agents API event stream contains only the lead’s turn, so for Sol we can see when the lead created, messaged, and waited on subagents, but not what it told them or what they did. Sol’s costs assume that its reported token usage includes the subagents.

Sol’s costs apply OpenAI’s base-tier list prices ($2 per million input tokens, $0.20 cached, $10 output) to reported token usage. Requests above the 272K-token long-context threshold are billed higher, which we could not evaluate because usage is reported per run rather than per request. Opus’s costs are the list-price totals Claude Code reports.

Because each setup ran once, the intervals and p-values in this post reflect variation across apps, not between repeated runs. The team runs also received the delegation instruction, so each comparison is between “subagents plus a delegation prompt” and neither. The two models ran through different harnesses, so comparisons across models mix the model with its harness.