Our first go at multi-agent: four run modes in one week
In one week in January we built a supervisor agent plus collaborative, Pipeline and self-organizing modes.
Once the single-agent loop had settled down at the end of December, we started on multi-agent. The motivation was simple: one agent working round by round is slow on long tasks; could we split the work across several?
Between January 6 and 13 we built four things.
1. A supervisor agent that watches and stays quiet
First came the supervisor. The need: once a long task is running, users want to know "where is it now", without the main agent interrupting itself to report.
It is just another consumer on the runtime event bus, alongside the desktop UI and mobile push:
- It subscribes only to low-frequency events: status changes, tool start and end, checkpoints, turn results. High-frequency text and thinking streams are ignored.
- Progress summaries reuse our existing action log first, instead of calling a model for every event.
- Output goes on its own channel; the CLI shows one status line, the desktop app has a separate panel.
- It never interjects but keeps control: it can pause or stop the main task at any time.
The acceptance bar was concrete: "Step 2 of 5 done, now: collecting dependency info" — it must name the current stage, not say "working hard on it".
2. Four modes
On top of the supervisor we built three ways for several agents to work together. With the existing single agent, four in total:
| Mode | Who decides | How work is split | Good for |
|---|---|---|---|
| single | One agent | Not split | Everyday development |
| collab | Main agent, at run time | Split as it goes, sub-agents on demand | Medium tasks that need flexibility |
| pipeline | Analyze, then plan | DAG generated up front, run layer by layer in parallel, optional multi-model review | Large, well-specified tasks |
| network | Coordinator + agents together | Broadcast the task, agents bid, negotiate roles | Exploratory long tasks |
collab is closest to what later became the familiar "sub-agent": the main agent decides whether to delegate, to whom, collects results and decides what's next. Flexible — but testing quickly showed the cost: the main agent spends lots of tokens on coordination and, after a few rounds of delegating, tends to lose the big picture.
pipeline is the opposite: a layer analyzes task complexity, complex tasks go to a planner that generates a DAG, and an executor runs it in parallel by dependency layer, optionally after several different models review the plan. Clear dependencies, traceable — but once the plan is set, it handles surprises poorly.
3. We had the name wrong
Funny thing: pipeline wasn't originally called pipeline. It was called network, and we kept calling it "self-organizing".
On January 10 we compared it with a paper on self-organizing agent networks and found it wasn't the same thing at all. In the paper, agents register their capabilities, bid on tasks they see, negotiate the split, review each other and replan during execution. Ours was a central planner producing a static DAG and executing it. That's a pipeline, not a self-organizing network.
So we renamed: the old network became pipeline, and the name network was kept for a real self-organizing mode, built from scratch.
4. The real self-organizing network, and its first-version problem
The new network mode followed the paper's layers: agent registration and capability discovery, task broadcast, bidding, negotiation, dynamic DAG, peer review. The UI got a card for each stage, so you could see who bid, how work was divided and how the reviews came out.
On January 11 and 12 we tested it hard and fixed a pile of issues — timeline rendering, context display, the final summary printing twice… but the fundamental problem was task decomposition.
The first version split the task by the number of agents that bid: three agents showed up, the task was cut into three. It sounds suitably decentralized and makes no sense — how many pieces a task needs depends on the task, not on how many hands go up.
We changed it so a model decomposes first: analyze the task, split it into subtasks, work out dependencies and execution layers; then match subtasks to existing agents, creating one on the spot if none fits. The DAG isn't fixed either — nodes can be added during execution. If decomposition fails, it falls back to the earlier hybrid strategy.
5. Context isolation in collab
collab got many fixes that week too, the most important being context isolation. Sub-agents initially shared the main agent's short-term memory and could see each other's intermediate steps. Sub-agents took the main agent's plan as their own task, and the main agent's context ballooned with sub-agents' tool output. Now each sub-agent has its own short-term memory and hands back only its result. We also fixed tool output being truncated inside sub-agents and the main agent's plan not being passed down.
Conclusion
That week we built and tested four modes. Testing showed that collab's main agent spent a lot of tokens on coordination, and every bidding and negotiation step in network was a model call; the first network version's habit of splitting tasks by bidder count only showed up in the logs; sub-agents sharing context interfered with each other until we gave each its own. We also found the mode we'd called "self-organizing" was really a pipeline, and renamed it.
For everyday coding, single is still what we use most. Which tasks multi-agent actually pays off on, that week didn't settle — it needs more real tasks.


