Skip to content
← All posts

Two sub-agents with the same name, one running 20 minutes with zero tool calls: eight root causes

A sub-agent ran 20 minutes with zero tool calls; eight problems were behind it.

A sub-agent ran 20 minutes with zero tool calls: duplicate dispatch, timeouts that never landed, wall-clock-only limits

On July 18 we were running a comparative benchmark: a refactoring task in a pet-raising project. Midway, the UI showed two sub-sessions with exactly the same name, "refactor pet growth". One ran a full 20 minutes with zero tool calls before a hard timeout killed it.

The score wasn't affected — the main agent finished the job itself. But it was worth digging, because it showed a lot was broken in the sub-agent layer, normally papered over by the main agent.

The scene

The stall log for that CLI run: 5 concurrent runs, 9 in-flight requests older than 20 seconds; one model request still pending after 719 seconds; several agent tool calls pending for two to three hundred seconds.

The sub-agent dispatch timeline:

TimeIDTaskOutcome
02:09:05Agent-1Read through the pet system—
02:15:00Agent-2Pet growth refactorMoved to background
02:15:37Agent-3Carry out growth refactorAborted after 480s
02:19:36Agent-4Fill in intimacy validationMoved to background
02:20:11Agent-5Implement intimacy fieldSame

Agent-3's prompt began "implement it directly in this repo, don't just give a plan". The main agent had read "moved to background" as "the sub-agent only gave a plan and didn't do anything", and pushed for a re-dispatch. It happened twice, 37 and 35 seconds apart.

Eight root causes

1. Going to the background returned a status, not a result. When a sub-agent didn't finish within a threshold it was moved to the background, and the main agent got only "moved to background automatically; you'll be notified when it's done". No result, no way to wait, no "don't dispatch this again". And the threshold was a parameter the model itself passed. The rational response was to dispatch another.

2. No task-level deduplication at any of three layers. Same-round batch execution split only into parallel and serial, with no dedup; the parallel-safe table explicitly listed agent, so two identical calls ran simultaneously; registration only checked whether the ID collided, not the task description. The dedup module's comment said "same-round dedup is handled by the parallel executor" — but that executor had long since been hollowed out into a types-only file. Nobody actually owned that responsibility.

A bug rode along: on an ID collision the new task was renamed Agent-1#2, #3, but later completion callbacks still used the original ID, so they landed on the first task and #3 stayed "running" until the 20-minute watchdog killed it. Session titles took the first 40 characters of the description; same description, two identical names.

3. Short timeouts silently didn't work. Every sub-agent type had a 3–8 minute max runtime. But the timeout aborted a signal and then kept awaiting, with no Promise.race. If execution was stuck somewhere that ignored the abort signal (a hung request, a dead stream), the await never ended. Only the 20-minute watchdog actually worked. "Killed after 20 minutes" wasn't a threshold set too high — the short timeout never took effect.

4. Kill decisions looked only at wall clock, not progress. The timeout check was "now minus start". Tool-call counts were only interpolated into a message, never used in the decision. The other guardrail tripped on more than 200K output tokens — the opposite direction, catching "too much output", never firing for a stuck agent with zero output. Every activity was timestamped, but nothing consumed it.

5. Model requests could hang forever. The stall monitor only logged — and after a few warnings it stopped logging. No abort, no retry, no escalation. No first-token timeout or idle timeout on the request path. The zero-tool sub-agent never even got its first token.

6. A global cap of 5, with foreground agents counting. Not per session; synchronous foreground sub-agents registered and counted too. Five runs shared one upstream and slowed each other down, turning cause 1's duplicate dispatches into an avalanche.

7. A killed sub-agent's output was thrown away entirely. Text already streamed was discarded on abort; the call that marked failure returned early because the status was already "aborted"; the main agent's completion notice carried the timeout message as its summary. Tokens burned, nothing recovered.

8. The tool description encouraged dispatching and never constrained it. Its first line was "send several in one message to run in parallel", with not a word about when not to. So work the main flow could finish in two minutes, like "fill in intimacy validation", got dispatched and took a dozen extra minutes. Claude Code's equivalent tool explicitly says "don't delegate a single lookup in a known file". Separately, the model parameter's description said "e.g. use a cheaper model for simple tasks", and the model itself downgraded the sub-agent a full generation.

What we fixed that day

ChangeStatus
Sub-agent execution uses Promise.race, with a hard 30-second backstop after a soft abortFixed
The background reply became an actionable instruction: "do not re-dispatch"Fixed
Dispatches identical after normalization (description or prompt) in the same session are blockedFixed
The tool description got a "when not to dispatch" barFixed
Zero-progress early stop: abort after 3 minutes with 0 tools and 0 output tokensFixed
Model downgrades only within the same generationFixed
Callbacks use the renamed IDFixed
Concurrency cap 5 → 3Fixed
First-token timeout + retryNot done — the real root of cause 5; zero-progress stop only shrinks the symptom from 20 minutes to 3
Record "last progress time" for idle detectionNot done
The stall monitor actually aborting or escalating past a thresholdNot done

A few trade-offs:

  • Dedup is deliberately conservative: only exact matches after normalization are blocked. Fuzzy matching would treat "edit views" and "edit routes" as duplicates and kill partitioned parallelism. The cost: the pair in this incident ("pet growth refactor" vs "carry out growth refactor") was worded differently and wouldn't be caught — that kind is handled by "do not re-dispatch" in the reply.
  • Zero progress means the strongest signal only: 0 tools and 0 output tokens. One emitted token and it doesn't count, to avoid killing legitimate long thinking. The cost: "did a little, then stuck" isn't caught.
  • Soft abort before the hard backstop: cooperative cancellation can recover partial output; force only if that fails.

What it meant for the benchmark

The desktop runs in that round took 7.9, 14.4 and 16.8 minutes — a wide spread, likely driven mostly by this idling rather than task difficulty. In other words, the time metric included both model ability and system idling.

Conclusion

Starting from the stall log and the dispatch timeline, we traced the idling to eight root causes: a background reply that made the model think nothing was done and re-dispatch, no dedup at three layers, short timeouts that never took effect, wall-clock-only kill decisions, model requests that could hang forever, a globally shared concurrency cap, output discarded on kill, and a tool description that only encouraged dispatching. We fixed eight items that day and left three, including the first-token timeout, for next. When analyzing timing data now, we first take out the part caused by system idling.

View source