One night of benchmarks, four P0s: npx hanging for 8 minutes, and two prompt lines worth 4.7x
One night of benchmarks found four P0s, each fixed in a few lines.
In mid-June we built our own benchmark — no toy problems, real projects:
- a React e-commerce project, about 300 files, for basic engineering tasks;
- an Express + Prisma + React full-stack chat app, 50 files and about 7,000 lines, for PR-sized cross-stack tasks;
- sub-agent stress tests: 5 and 8 explore sub-agents in parallel.
The harness launches the CLI, captures the full event stream and saves every request sent to the model. On the night of June 15 we ran it end to end for the first time, and the next day we had four P0s.
P0-B: one npx call hung for 8 minutes 20 seconds
Symptom: the agent ran npx tsc --noEmit, timed out after 8:20, and retried again and again.
Investigation: ps, lsof and sample on the stuck process: its stdin/stdout were sockets, and it was blocked waiting on IO. The same command in a terminal exited in 0.3 seconds.
Root cause: the shell command was wrapped twice. We already built zsh -lc "<command>", then passed it to the process library with shell: true, which wraps it again in /bin/sh -c. The result:
/bin/sh -c "zsh -lc npx tsc --noEmit; ..."After sh parsed it, zsh -c received only npx; tsc --noEmit became zsh positional arguments that npx never saw. npx with no arguments enters interactive mode and waits on stdin — a socket that never sends anything. Deadlock.
Fix: drop shell: true and explicitly ignore stdin so tools like npx can't wait for input. 8:20 → 5.5 seconds, 90x.
The scary part was the latency: the commit that introduced it had added a fallback execution path the day before, and the benchmark caught it the next day. Without the benchmark, every user would have waited 8 minutes on every npx tsc, with no idea why.
P0-D: a fake pass
Symptom: a task to "add coupons to the shop" ran 65 seconds, 14 rounds, 22 tool calls, zero edits, zero file writes, was marked passed, and the agent's final answer was "the coupon system is essentially complete, all 130 tests pass".
A PR-sized task touching 8 files can't be done without a single edit.
Root cause: the harness didn't reset the project between runs. The 8 coupon files written by the previous run were still there; the agent explored, found "already implemented", and honestly reported done.
Fix: git reset --hard + git clean at the start of every run. Also a to-do: the pass check should look at git diff too — if a build-something task produces an empty diff, mark it a fake pass.
That's a harness bug rather than a product bug, but it matters: with a dirty environment, a conclusion like "this setup is faster and better" can simply be wrong.
P0-1: 40 rounds fixing a function that didn't exist
Symptom: a bug-hunt task said updateQuantity(id, n) had a ghost-item bug. The cart code had no `updateQuantity` method at all.
The agent's trajectory: guess → write and test → "4 pass, 1 is my test's fault" → explore "high concurrency" → "extreme repetition"… only in round 41 did it grep for updateQuantity, and it ended up fixing an unrelated increase function. 913.9 seconds, 40 rounds, 46 tool calls.
Fix: prompt only. Two lines added to the "engineering principles":
- Verify the task's premises first. If the task names a specific function, file path, method or class, grep to confirm it exists. If not, ask "I can't find X — did you mean Y?" Don't guess.
- Step back after 3 failures in the same direction. List the failed hypotheses, then ask or change angle.
The same task after the fix: one run took 3 rounds, 5 tool calls, 193.8 seconds — read the code and honestly said "this function doesn't exist", with zero changes to the project; another read the code and pinpointed the real bug (addItem used a stale stock value), with a one-line fix and a regression test. 4.7x faster, 13x fewer rounds, 9x fewer tool calls.
No architecture change, no tool change — two lines of prompt.
P0-2: five "parallel" sub-agents in a queue
Symptom: we asked the model for "five explore sub-agents scanning five directories in parallel". It did emit five explore calls in the same millisecond, but the sub-agents' first requests were spaced 16, 7, 11 and 18 seconds apart — each started after the previous finished. Five in parallel took 108 seconds; eight took 106. Nearly identical.
Root cause: the allowlist of parallel-safe tools included agent but not `explore`. Both tools share the same logic — separate session, separate runner, read-only tools — and are fully parallel-safe. But the scheduler treated explore as state-changing and awaited each in turn.
Fix: one line added to the allowlist. 108.1 → 57.1 seconds; excluding the main agent's own 14 seconds, the sub-agent part went from 94 to 43.
It also shows why, two months earlier, we moved concurrency decisions from an allowlist to "each tool decides from its arguments" — one missing entry is a P0, and nothing errors.
Real-task sanity check
After the fixes we ran two PR-sized tasks on the full-stack app:
| Task | Time | Rounds | Tools | Output | Verified |
|---|---|---|---|---|---|
| Fix 12 lint errors | 249s | 31 | 37 | 6 files changed | npm run lint passes |
| Cross-stack "favorite messages" | 355s | 64 | 93 | 12 files (4 new, 8 changed) | prisma + build pass, no new lint errors |
One detail in the second task pleased us: the agent separated "lint errors the project already had" from "ones introduced this time", neither blaming itself for the old ones nor fixing them on the side.
Conclusion
That night's benchmark found and fixed four P0s: a doubly wrapped shell hanging npx for 8 minutes (90x after the fix), a harness that didn't reset producing fake passes, the agent guessing at a nonexistent function (4.7x after two prompt lines), and parallel explores running serially because the allowlist missed them (1.9x after the fix). The npx bug had been introduced only the day before and the benchmark caught it the next day.
We also found single runs were very noisy — two setups that looked 4x apart could just be random variation — so comparisons since then use at least 3 runs. From that day on, before changing prompts or the runner, we run this benchmark first, with red lines on key tasks: the bug hunt over 300 seconds or 10 rounds, five parallel explores over 80 seconds, a simple shell command over 60 seconds, or a build task marked passed with an empty diff — any one of them means investigate.


