28 experiments on multi-agent clusters: speed is set by the serial part, and over-constraining makes a cluster slower than one agent
28 experiments comparing single agents with clusters: cluster speed is set by the serial part.
Our first multi-agent attempt in January left one question open: "when is multi-agent actually worth it?" At the end of May, preparing to build a cluster agent framework (internally called Colony), we decided not to write the framework first — answer the question with data first.
The goal was blunt: can N agents in parallel finish in 1 hour what takes a human developer 8, without ever being slower because agents get in each other's way?
How we ran it
Everything ran on our real gateway and the real Neox agent kernel, all with DeepSeek V4 Flash. We wrote seven evaluation suites:
- Compounding curve: one agent builds N modules in a row; how does time per module change?
- Contention: agents in parallel with a shared workspace vs isolated workspaces;
- Multi-process end to end: N modules, K agents, real parallel processes;
- One vs many, heterogeneous large tasks, dynamic refill;
- A role-based real app: "architect / backend / frontend" building an app with 59 assertions.
28 groups in total.
Eleven laws
| # | Law | Measured |
|---|---|---|
| L1 | The gateway isn't the bottleneck | Zero slowdown under cross-process concurrency — stop blaming rate limits |
| L2 | Parallelism is real | Separate processes + isolated workspaces = zero contention; a shared workspace is 2.6x slower |
| L3 | Context compounding is the moat | A single agent took 27.5s for module 2 and 6.3s by module 12 — 4.33x faster, driven by cache hits |
| L4 | Compounding keeps deepening as the project grows | 4.9s per module at 30 modules |
| L5 | The crossover is around 20 modules | With 12 identical modules, multi-agent ran at 0.73x of single; at 30, 1.23x |
| L6 | Cold starts inflate work 3.2x | A new session per task means 3.2x the total work of a compounding single agent, wiping out parallel gains |
| L7 | Wall time is set by the slowest indivisible task | One hard, high-variance atomic task (99–274s) gated the whole group |
| L8 | Coordination can cost more than the task | A heavy architect + contract phase took 70% of cluster wall time — 2x slower than one agent building the whole app |
| L9 | Task shape decides the speedup | Small and uniform: capped at 1.3x; large, heterogeneous, independent: 2.6–3.9x in theory |
| L10 | Quality is always equal | Integration tests matched in every experiment — the model sets quality, orchestration doesn't |
| L11 | Clusters cost 3–6x the tokens | Architect phase plus per-agent startup overhead — you're buying speed with money |
L3 and L8 surprised us most.
L3: a single agent gets faster as it goes. Building modules one after another in the same session, the code it read and the conventions it set stay in context, mostly as cache hits. So for multi-agent to be faster, it first has to offset that compounding.
L8: coordinating can cost more than working. We first designed it as "an architect fixes the interface contracts, then hands pieces to agents", which felt safest. That contract phase took 70% of cluster wall time, and overall speed was 0.35x of a single agent.
Four structures, head to head
On June 1 we ran one app (59 assertions) under four organizational structures:
| Rank | Structure | Median wall time | Serial share |
|---|---|---|---|
| 1 | Single agent | 71.9s | 0% |
| 2 | Mission command (interfaces given up front, everything parallel) | 113.2s | ~0% |
| 3 | Orchestrator-worker (light lead plans, workers execute) | 119.6s | 20% |
| 4 | Bureaucracy (heavy architect fixes contracts first) | 140.4s | 68% |
The ranking matches the serial share exactly. That's Amdahl's law: cluster speed is set by the serial part. Removing the heavy serial design phase is the first lever; mission command was 1.24x faster than bureaucracy.
This app wasn't big enough, so the single agent still came first. But it shows what a cluster should look like.
Read against organizational theory
At this point we went through organizational theory and found a ready-made counterpart for nearly every law:
- Anthropic vs Cognition. Two opposing write-ups from 2025: Anthropic's research system used "a lead agent plans + 3–5 sub-agents in parallel" and scored 90.2% above a single agent internally; Cognition wrote "Don't Build Multi-Agents", with the example of building Flappy Bird where one sub-agent made a Mario-style background and another drew a bird in a clashing style — the implicit decision "match the original art" was lost in the split. Both are right: reading tasks (research, retrieval) parallelize well; writing tasks (coding) don't, because writing needs globally consistent decisions.
- Smith's pin factory. Division of labor raised output per worker 240x — given a conveyor that made handoffs free. Agents have no free conveyor, so split coarsely: by subsystem, not by function. We once cut work into tiny functions of a few dozen seconds each, and startup overhead swamped the output.
- Mission command (Auftragstaktik). The Prussian approach: commanders give intent and objectives, not detailed orders. That's why "mission command" won — each agent gets "intent + acceptance criteria + interfaces + why", and nothing more.
- Coase's theory of the firm. A firm should stop growing where internal coordination cost equals marginal output. Same for clusters: more agents isn't always better.
- Brooks's law. Communication paths grow as N², newcomers need ramp-up, and an indivisible critical path doesn't speed up with more people. These map to our startup overhead, "don't do N² messaging", and L7's bottleneck task.
- Amazon's API mandate. Teams may only interact through interfaces; no shared databases. That's L2: a shared workspace is 2.6x slower; isolated, there's no contention.
- The learning curve. Veterans are faster than newcomers; organizations keep knowledge in institutional memory. That's L3 and L6: cold-starting a fresh agent per task is "fire and rehire" every time, losing all institutional knowledge.
- Theory of constraints. System throughput equals the bottleneck's throughput. That's L7: don't spread effort evenly — give the hardest task the strongest model or split it further, and don't let other agents sit idle waiting.
Where we had gone wrong
Seen this way, almost every choice in our earlier design was "restraining the agent":
| What we did | What we should do | |
|---|---|---|
| Decisions | A heavy architect pins every detail | Give intent, interfaces and acceptance criteria only |
| Split size | Tiny functions of a few dozen seconds | By domain and subsystem |
| Shared state | A shared workspace | A workspace per agent; interact only through interfaces and tests |
| Coordination | Heavy up-front design documents | Interfaces + integration tests; coordinate through artifacts, not messages |
| Agent form | Cold start per task | Persistent-session "veterans" + shared memory |
| Bottleneck task | Even effort | Targeted resources |
| When to split | Always | Only large, heterogeneous, independent tasks |
Conclusion
From the 28 experiments and the four-structure comparison we concluded: on small tasks a single agent is faster thanks to context compounding, with the crossover around 20 uniform modules; cluster speed is set by the serial part, and a heavy up-front design phase makes a cluster slower than one agent; cold-started sub-agents inflate total work 3.2x; clusters don't improve quality — they buy speed at 3–6x the tokens.
So every further step of Colony first has to show, on the same task, that it is no worse than a single agent and genuinely faster on the parallel part; otherwise we don't continue.


