---
title: "Two sub-agents with the same name, one running 20 minutes with zero tool calls: eight root causes"
date: "2026-07-18"
tag: "Engineering"
lang: "en"
reading_minutes: 9
source: "https://neox-dev.com/blog/subagent-idle-20-minutes-zero-tools"
alternate: "https://neox-dev.com/md/blog/subagent-idle-20-minutes-zero-tools.md?lang=zh"
---

# Two sub-agents with the same name, one running 20 minutes with zero tool calls: eight root causes

> Running a benchmark on July 18, the UI showed two identical sub-sessions, "refactor pet growth"; one ran a full 20 minutes with zero tool calls before a hard timeout killed it. Digging in, we found eight root causes: a background-status reply the model read as "did nothing" and re-dispatched, no deduplication at any of three layers, 3–8 minute timeouts that never took effect, wall-clock-only kill decisions, output thrown away on kill…

On July 18 we were running a comparative benchmark: a refactoring task in a pet-raising project. Midway, the UI showed two sub-sessions with **exactly the same name**, "refactor pet growth". One ran a full 20 minutes with **zero tool calls** before a hard timeout killed it.

The score wasn't affected — the main agent finished the job itself. But it was worth digging, because it showed a lot was broken in the sub-agent layer, normally papered over by the main agent.

## The scene

The stall log for that CLI run: 5 concurrent runs, 9 in-flight requests older than 20 seconds; one model request still pending after **719 seconds**; several `agent` tool calls pending for two to three hundred seconds.

The sub-agent dispatch timeline:

| Time | ID | Task | Outcome |
|---|---|---|---|
| 02:09:05 | Agent-1 | Read through the pet system | — |
| 02:15:00 | Agent-2 | Pet growth refactor | Moved to background |
| 02:15:37 | Agent-3 | **Carry out** growth refactor | Aborted after 480s |
| 02:19:36 | Agent-4 | Fill in intimacy validation | Moved to background |
| 02:20:11 | Agent-5 | **Implement** intimacy field | Same |

Agent-3's prompt began "implement it directly in this repo, **don't just give a plan**". The main agent had read "moved to background" as "the sub-agent only gave a plan and didn't do anything", and pushed for a re-dispatch. It happened twice, 37 and 35 seconds apart.

## Eight root causes

**1. Going to the background returned a status, not a result.** When a sub-agent didn't finish within a threshold it was moved to the background, and the main agent got only "moved to background automatically; you'll be notified when it's done". No result, no way to wait, no "don't dispatch this again". And the threshold was a parameter the model itself passed. The rational response was to dispatch another.

**2. No task-level deduplication at any of three layers.** Same-round batch execution split only into parallel and serial, with no dedup; the parallel-safe table explicitly listed `agent`, so two identical calls ran simultaneously; registration only checked whether the ID collided, not the task description. The dedup module's comment said "same-round dedup is handled by the parallel executor" — but that executor had long since been hollowed out into a types-only file. **Nobody actually owned that responsibility.**

A bug rode along: on an ID collision the new task was renamed `Agent-1#2`, `#3`, but later completion callbacks still used the original ID, so they landed on the first task and `#3` stayed "running" until the 20-minute watchdog killed it. Session titles took the first 40 characters of the description; same description, two identical names.

**3. Short timeouts silently didn't work.** Every sub-agent type had a 3–8 minute max runtime. But the timeout aborted a signal and then **kept awaiting**, with no `Promise.race`. If execution was stuck somewhere that ignored the abort signal (a hung request, a dead stream), the await never ended. Only the 20-minute watchdog actually worked. **"Killed after 20 minutes" wasn't a threshold set too high — the short timeout never took effect.**

**4. Kill decisions looked only at wall clock, not progress.** The timeout check was "now minus start". Tool-call counts were only interpolated into a message, never used in the decision. The other guardrail tripped on more than 200K output tokens — the opposite direction, catching "too much output", never firing for a stuck agent with zero output. Every activity was timestamped, but nothing consumed it.

**5. Model requests could hang forever.** The stall monitor only logged — and after a few warnings it stopped logging. No abort, no retry, no escalation. No first-token timeout or idle timeout on the request path. The zero-tool sub-agent never even got its first token.

**6. A global cap of 5, with foreground agents counting.** Not per session; synchronous foreground sub-agents registered and counted too. Five runs shared one upstream and slowed each other down, turning cause 1's duplicate dispatches into an avalanche.

**7. A killed sub-agent's output was thrown away entirely.** Text already streamed was discarded on abort; the call that marked failure returned early because the status was already "aborted"; the main agent's completion notice carried the timeout message as its summary. **Tokens burned, nothing recovered.**

**8. The tool description encouraged dispatching and never constrained it.** Its first line was "send several in one message to run in parallel", with not a word about when not to. So work the main flow could finish in two minutes, like "fill in intimacy validation", got dispatched and took a dozen extra minutes. Claude Code's equivalent tool explicitly says "don't delegate a single lookup in a known file". Separately, the `model` parameter's description said "e.g. use a cheaper model for simple tasks", and the model itself downgraded the sub-agent a full generation.

## What we fixed that day

| Change | Status |
|---|---|
| Sub-agent execution uses `Promise.race`, with a hard 30-second backstop after a soft abort | Fixed |
| The background reply became an actionable instruction: "do not re-dispatch" | Fixed |
| Dispatches identical after normalization (description or prompt) in the same session are blocked | Fixed |
| The tool description got a "when not to dispatch" bar | Fixed |
| Zero-progress early stop: abort after 3 minutes with 0 tools and 0 output tokens | Fixed |
| Model downgrades only within the same generation | Fixed |
| Callbacks use the renamed ID | Fixed |
| Concurrency cap 5 → 3 | Fixed |
| First-token timeout + retry | Not done — the real root of cause 5; zero-progress stop only shrinks the symptom from 20 minutes to 3 |
| Record "last progress time" for idle detection | Not done |
| The stall monitor actually aborting or escalating past a threshold | Not done |

A few trade-offs:

- **Dedup is deliberately conservative**: only exact matches after normalization are blocked. Fuzzy matching would treat "edit views" and "edit routes" as duplicates and kill partitioned parallelism. The cost: the pair in this incident ("pet growth refactor" vs "carry out growth refactor") was worded differently and wouldn't be caught — that kind is handled by "do not re-dispatch" in the reply.
- **Zero progress means the strongest signal only**: 0 tools **and** 0 output tokens. One emitted token and it doesn't count, to avoid killing legitimate long thinking. The cost: "did a little, then stuck" isn't caught.
- **Soft abort before the hard backstop**: cooperative cancellation can recover partial output; force only if that fails.

## What it meant for the benchmark

The desktop runs in that round took 7.9, 14.4 and 16.8 minutes — a wide spread, likely driven mostly by this idling rather than task difficulty. In other words, the time metric included both model ability and system idling.

## Conclusion

Starting from the stall log and the dispatch timeline, we traced the idling to eight root causes: a background reply that made the model think nothing was done and re-dispatch, no dedup at three layers, short timeouts that never took effect, wall-clock-only kill decisions, model requests that could hang forever, a globally shared concurrency cap, output discarded on kill, and a tool description that only encouraged dispatching. We fixed eight items that day and left three, including the first-token timeout, for next. When analyzing timing data now, we first take out the part caused by system idling.
