---
title: "One night of benchmarks, four P0s: npx hanging for 8 minutes, and two prompt lines worth 4.7x"
date: "2026-06-16"
tag: "Engineering"
lang: "en"
reading_minutes: 8
source: "https://neox-dev.com/blog/one-night-bench-four-p0"
alternate: "https://neox-dev.com/md/blog/one-night-bench-four-p0.md?lang=zh"
---

# One night of benchmarks, four P0s: npx hanging for 8 minutes, and two prompt lines worth 4.7x

> In mid-June we built a benchmark running real tasks against a 300-file React shop and a 7,000-line full-stack app. One night caught four P0s: a doubly wrapped shell that made npx hang forever; a harness that didn't reset between runs, producing "fake passes"; an agent spending 40 rounds fixing a function that didn't exist; and five "parallel" explore sub-agents that were actually queuing. Each fix was a few lines.

In mid-June we built our own benchmark — no toy problems, real projects:

- a React e-commerce project, about 300 files, for basic engineering tasks;
- an Express + Prisma + React full-stack chat app, 50 files and about 7,000 lines, for PR-sized cross-stack tasks;
- sub-agent stress tests: 5 and 8 explore sub-agents in parallel.

The harness launches the CLI, captures the full event stream and saves every request sent to the model. On the night of June 15 we ran it end to end for the first time, and the next day we had four P0s.

## P0-B: one npx call hung for 8 minutes 20 seconds

**Symptom:** the agent ran `npx tsc --noEmit`, timed out after 8:20, and retried again and again.

**Investigation:** `ps`, `lsof` and `sample` on the stuck process: its stdin/stdout were sockets, and it was blocked waiting on IO. The same command in a terminal exited in 0.3 seconds.

**Root cause:** the shell command was wrapped twice. We already built `zsh -lc "<command>"`, then passed it to the process library with `shell: true`, which wraps it again in `/bin/sh -c`. The result:

```
/bin/sh -c "zsh -lc npx tsc --noEmit; ..."
```

After `sh` parsed it, `zsh -c` received only `npx`; `tsc --noEmit` became zsh positional arguments that npx never saw. npx with no arguments enters interactive mode and waits on stdin — a socket that never sends anything. Deadlock.

**Fix:** drop `shell: true` and explicitly ignore stdin so tools like npx can't wait for input. **8:20 → 5.5 seconds, 90x.**

The scary part was the latency: the commit that introduced it had added a fallback execution path the day before, and **the benchmark caught it the next day.** Without the benchmark, every user would have waited 8 minutes on every `npx tsc`, with no idea why.

## P0-D: a fake pass

**Symptom:** a task to "add coupons to the shop" ran 65 seconds, 14 rounds, 22 tool calls, **zero edits, zero file writes**, was marked passed, and the agent's final answer was "the coupon system is essentially complete, all 130 tests pass".

A PR-sized task touching 8 files can't be done without a single edit.

**Root cause:** the harness didn't reset the project between runs. The 8 coupon files written by the previous run were still there; the agent explored, found "already implemented", and honestly reported done.

**Fix:** `git reset --hard` + `git clean` at the start of every run. Also a to-do: the pass check should look at `git diff` too — if a build-something task produces an empty diff, mark it a fake pass.

That's a harness bug rather than a product bug, but it matters: with a dirty environment, a conclusion like "this setup is faster and better" can simply be wrong.

## P0-1: 40 rounds fixing a function that didn't exist

**Symptom:** a bug-hunt task said `updateQuantity(id, n)` had a ghost-item bug. The cart code **had no `updateQuantity` method at all.**

The agent's trajectory: guess → write and test → "4 pass, 1 is my test's fault" → explore "high concurrency" → "extreme repetition"… only in round 41 did it grep for `updateQuantity`, and it ended up fixing an unrelated `increase` function. **913.9 seconds, 40 rounds, 46 tool calls.**

**Fix:** prompt only. Two lines added to the "engineering principles":

1. **Verify the task's premises first.** If the task names a specific function, file path, method or class, grep to confirm it exists. If not, ask "I can't find X — did you mean Y?" Don't guess.
2. **Step back after 3 failures in the same direction.** List the failed hypotheses, then ask or change angle.

**The same task after the fix:** one run took 3 rounds, 5 tool calls, 193.8 seconds — read the code and honestly said "this function doesn't exist", with zero changes to the project; another read the code and pinpointed the real bug (`addItem` used a stale stock value), with a one-line fix and a regression test. **4.7x faster, 13x fewer rounds, 9x fewer tool calls.**

No architecture change, no tool change — two lines of prompt.

## P0-2: five "parallel" sub-agents in a queue

**Symptom:** we asked the model for "five explore sub-agents scanning five directories in parallel". It did emit five `explore` calls in the same millisecond, but the sub-agents' first requests were spaced 16, 7, 11 and 18 seconds apart — each started after the previous finished. Five in parallel took 108 seconds; eight took 106. Nearly identical.

**Root cause:** the allowlist of parallel-safe tools included `agent` but **not `explore`.** Both tools share the same logic — separate session, separate runner, read-only tools — and are fully parallel-safe. But the scheduler treated `explore` as state-changing and awaited each in turn.

**Fix:** one line added to the allowlist. **108.1 → 57.1 seconds**; excluding the main agent's own 14 seconds, the sub-agent part went from 94 to 43.

It also shows why, two months earlier, we moved concurrency decisions from an allowlist to "each tool decides from its arguments" — one missing entry is a P0, and nothing errors.

## Real-task sanity check

After the fixes we ran two PR-sized tasks on the full-stack app:

| Task | Time | Rounds | Tools | Output | Verified |
|---|---|---|---|---|---|
| Fix 12 lint errors | 249s | 31 | 37 | 6 files changed | `npm run lint` passes |
| Cross-stack "favorite messages" | 355s | 64 | 93 | 12 files (4 new, 8 changed) | prisma + build pass, no new lint errors |

One detail in the second task pleased us: the agent separated "lint errors the project already had" from "ones introduced this time", neither blaming itself for the old ones nor fixing them on the side.

## Conclusion

That night's benchmark found and fixed four P0s: a doubly wrapped shell hanging npx for 8 minutes (90x after the fix), a harness that didn't reset producing fake passes, the agent guessing at a nonexistent function (4.7x after two prompt lines), and parallel explores running serially because the allowlist missed them (1.9x after the fix). The npx bug had been introduced only the day before and the benchmark caught it the next day.

We also found single runs were very noisy — two setups that looked 4x apart could just be random variation — so comparisons since then use at least 3 runs. From that day on, before changing prompts or the runner, we run this benchmark first, with red lines on key tasks: the bug hunt over 300 seconds or 10 rounds, five parallel explores over 80 seconds, a simple shell command over 60 seconds, or a build task marked passed with an empty diff — any one of them means investigate.
