---
title: "28 experiments on multi-agent clusters: speed is set by the serial part, and over-constraining makes a cluster slower than one agent"
date: "2026-06-01"
tag: "Engineering"
lang: "en"
reading_minutes: 11
source: "https://neox-dev.com/blog/colony-28-experiments-cluster-speed"
alternate: "https://neox-dev.com/md/blog/colony-28-experiments-cluster-speed.md?lang=zh"
---

# 28 experiments on multi-agent clusters: speed is set by the serial part, and over-constraining makes a cluster slower than one agent

> At the end of May we set out to build a cluster agent framework to turn 8 hours of work into 1. After 28 experiment groups the conclusions were far from what we expected: on small tasks a single agent wins through context compounding; a heavy "architect" phase can take 70% of cluster time; cold-started sub-agents inflate total work 3.2x. Here are the laws and the data, read against organizational theory — pin factories, the Prussian army and Amazon hit many of the same walls long ago.

Our first multi-agent attempt in January left one question open: "when is multi-agent actually worth it?" At the end of May, preparing to build a cluster agent framework (internally called Colony), we decided not to write the framework first — **answer the question with data first.**

The goal was blunt: can N agents in parallel finish in 1 hour what takes a human developer 8, without ever being slower because agents get in each other's way?

## How we ran it

Everything ran on our real gateway and the real Neox agent kernel, all with DeepSeek V4 Flash. We wrote seven evaluation suites:

- **Compounding curve**: one agent builds N modules in a row; how does time per module change?
- **Contention**: agents in parallel with a shared workspace vs isolated workspaces;
- **Multi-process end to end**: N modules, K agents, real parallel processes;
- **One vs many**, **heterogeneous large tasks**, **dynamic refill**;
- **A role-based real app**: "architect / backend / frontend" building an app with 59 assertions.

28 groups in total.

## Eleven laws

| # | Law | Measured |
|---|---|---|
| L1 | The gateway isn't the bottleneck | Zero slowdown under cross-process concurrency — stop blaming rate limits |
| L2 | Parallelism is real | Separate processes + isolated workspaces = zero contention; a shared workspace is 2.6x slower |
| L3 | Context compounding is the moat | A single agent took 27.5s for module 2 and 6.3s by module 12 — 4.33x faster, driven by cache hits |
| L4 | Compounding keeps deepening as the project grows | 4.9s per module at 30 modules |
| L5 | The crossover is around 20 modules | With 12 identical modules, multi-agent ran at 0.73x of single; at 30, 1.23x |
| L6 | Cold starts inflate work 3.2x | A new session per task means 3.2x the total work of a compounding single agent, wiping out parallel gains |
| L7 | Wall time is set by the slowest indivisible task | One hard, high-variance atomic task (99–274s) gated the whole group |
| L8 | Coordination can cost more than the task | A heavy architect + contract phase took 70% of cluster wall time — 2x slower than one agent building the whole app |
| L9 | Task shape decides the speedup | Small and uniform: capped at 1.3x; large, heterogeneous, independent: 2.6–3.9x in theory |
| L10 | Quality is always equal | Integration tests matched in every experiment — the model sets quality, orchestration doesn't |
| L11 | Clusters cost 3–6x the tokens | Architect phase plus per-agent startup overhead — you're buying speed with money |

L3 and L8 surprised us most.

**L3: a single agent gets faster as it goes.** Building modules one after another in the same session, the code it read and the conventions it set stay in context, mostly as cache hits. So for multi-agent to be faster, it first has to offset that compounding.

**L8: coordinating can cost more than working.** We first designed it as "an architect fixes the interface contracts, then hands pieces to agents", which felt safest. That contract phase took 70% of cluster wall time, and overall speed was 0.35x of a single agent.

## Four structures, head to head

On June 1 we ran one app (59 assertions) under four organizational structures:

| Rank | Structure | Median wall time | Serial share |
|---|---|---|---|
| 1 | Single agent | 71.9s | 0% |
| 2 | Mission command (interfaces given up front, everything parallel) | 113.2s | ~0% |
| 3 | Orchestrator-worker (light lead plans, workers execute) | 119.6s | 20% |
| 4 | Bureaucracy (heavy architect fixes contracts first) | 140.4s | 68% |

The ranking matches the serial share exactly. That's Amdahl's law: **cluster speed is set by the serial part.** Removing the heavy serial design phase is the first lever; mission command was 1.24x faster than bureaucracy.

This app wasn't big enough, so the single agent still came first. But it shows what a cluster should look like.

## Read against organizational theory

At this point we went through organizational theory and found a ready-made counterpart for nearly every law:

- **Anthropic vs Cognition.** Two opposing write-ups from 2025: Anthropic's research system used "a lead agent plans + 3–5 sub-agents in parallel" and scored 90.2% above a single agent internally; Cognition wrote "Don't Build Multi-Agents", with the example of building Flappy Bird where one sub-agent made a Mario-style background and another drew a bird in a clashing style — the implicit decision "match the original art" was lost in the split. Both are right: **reading tasks (research, retrieval) parallelize well; writing tasks (coding) don't,** because writing needs globally consistent decisions.
- **Smith's pin factory.** Division of labor raised output per worker 240x — given a conveyor that made handoffs free. Agents have no free conveyor, so split coarsely: by subsystem, not by function. We once cut work into tiny functions of a few dozen seconds each, and startup overhead swamped the output.
- **Mission command (Auftragstaktik).** The Prussian approach: commanders give intent and objectives, not detailed orders. That's why "mission command" won — each agent gets "intent + acceptance criteria + interfaces + why", and nothing more.
- **Coase's theory of the firm.** A firm should stop growing where internal coordination cost equals marginal output. Same for clusters: more agents isn't always better.
- **Brooks's law.** Communication paths grow as N², newcomers need ramp-up, and an indivisible critical path doesn't speed up with more people. These map to our startup overhead, "don't do N² messaging", and L7's bottleneck task.
- **Amazon's API mandate.** Teams may only interact through interfaces; no shared databases. That's L2: a shared workspace is 2.6x slower; isolated, there's no contention.
- **The learning curve.** Veterans are faster than newcomers; organizations keep knowledge in institutional memory. That's L3 and L6: cold-starting a fresh agent per task is "fire and rehire" every time, losing all institutional knowledge.
- **Theory of constraints.** System throughput equals the bottleneck's throughput. That's L7: don't spread effort evenly — give the hardest task the strongest model or split it further, and don't let other agents sit idle waiting.

## Where we had gone wrong

Seen this way, almost every choice in our earlier design was "restraining the agent":

| | What we did | What we should do |
|---|---|---|
| Decisions | A heavy architect pins every detail | Give intent, interfaces and acceptance criteria only |
| Split size | Tiny functions of a few dozen seconds | By domain and subsystem |
| Shared state | A shared workspace | A workspace per agent; interact only through interfaces and tests |
| Coordination | Heavy up-front design documents | Interfaces + integration tests; coordinate through artifacts, not messages |
| Agent form | Cold start per task | Persistent-session "veterans" + shared memory |
| Bottleneck task | Even effort | Targeted resources |
| When to split | Always | Only large, heterogeneous, independent tasks |

## Conclusion

From the 28 experiments and the four-structure comparison we concluded: on small tasks a single agent is faster thanks to context compounding, with the crossover around 20 uniform modules; cluster speed is set by the serial part, and a heavy up-front design phase makes a cluster slower than one agent; cold-started sub-agents inflate total work 3.2x; clusters don't improve quality — they buy speed at 3–6x the tokens.

So every further step of Colony first has to show, on the same task, that it is no worse than a single agent and genuinely faster on the parallel part; otherwise we don't continue.
