Our benchmark was punishing "ask first": a mistake in how we scored
Claude Code stopped to ask about an ambiguity and scored the baseline; the scoring was wrong.
In July we were running an agent comparison: same model, same task, handed unattended to Neox, Claude Code and Codex, scored by acceptance tests.
In the Opus 4.8 group, Claude Code's three runs were identical: 1.3 minutes, 0 files, 0 lines changed, 11/18. And 11/18 is exactly the "do nothing" baseline.
It didn't fail
We ruled out the obvious first: it ended normally, with no permission denials; it did call tools — 4 file reads, 7 shell commands, 4 test runs; and it used the specified model.
Reproducing in an isolated copy, we got its full reply. In essence:
1. Achievement badges already exist and are fully implemented. There are 14 badges, a separate table with a unique constraint, unlock logic wired into 4 routes, and the detail page already renders the badge grid. Item 3 is done, and the achievement part of item 4 is done too. I'd rather not rebuild it. 2. There is no "intimacy" stat. The four stats are happiness, hunger, energy and health. Three of your four match; the fourth is health, not intimacy. This is a fork you need to decide: (a) you misremembered the field name — treat it as health, no schema change, contained scope; (b) you really want a fifth field — that needs a schema migration and changes to all 6 update call sites. Much more work.
It also reported a real bug along the way: the streak calculation ran a dead LIMIT 0 query, discarded the result, and recomputed with a second query.
In unattended mode nobody could answer, so it printed its findings and ended normally. Zero files.
Our task description really was wrong in two places
| What we wrote | What the code actually had |
|---|---|
| "Add achievement badges (at least 3)" | 14 already existed, fully implemented and on screen |
| "Pet stats (hunger/happiness/energy/intimacy)" | The fourth field was health; there was no intimacy |
The real finding: different strategies for ambiguity
| Agent | On ambiguity | Result under this scoring |
|---|---|---|
| Claude Code | Stopped and laid out the fork for the user | 11/18 (= baseline) |
| Neox | Decided itself: the prompt it gave a sub-agent said "if the health field currently stands for health broadly, to avoid breaking compatibility, add an intimacy field and migrate" | 17–18/18 |
| Codex | Decided itself and went ahead | 17–18/18 |
Neox noticed the same ambiguity; it just made the call itself. So this isn't a gap in "who understood the code" — it's a difference in how ambiguity is handled. Unattended, "asking" means zero output.
The scoring was wrong
Our scoring rewarded charging ahead and punished careful clarification. Writing Claude Code's 11/18 into a report as "weaker capability" would be a false conclusion.
What we did:
- This batch's 11/18 for Claude Code is not presented as a capability score but described separately as "asked for clarification due to ambiguity; no output when unattended" — unscorable under this rubric, not a low score;
- Added an unattended clause to the task prompt: "This is an unattended automated task; nobody can answer your questions. If the requirements conflict with the existing code or are ambiguous, make a reasonable judgment, explain your choice in your final reply, and carry on — don't stop to wait for confirmation";
- Reran Claude Code three times with the new prompt for a scorable number;
- Reported both results: "it chose to ask" under the original prompt, "its actual ability" under the new one. The former is itself valuable evidence of a behavioral difference.
Should we fix the task description?
We decided not to:
- The bulk of the acceptance tests doesn't depend on the two errors, so the actual new work is unaffected;
- "Achievements already exist" applies equally to every agent, so it isn't unfair;
- The "intimacy or health" ambiguity is itself a valuable test point — it separates "noticed and handled the ambiguity" from "never saw it";
- Changing the prompt would break comparability with the other model groups.
The unattended clause is enough: it standardizes "what to do once you notice an ambiguity" without removing the ambiguity.
Conclusion
Reproducing Claude Code's full reply, we found its 11/18 wasn't a capability problem: having noticed the task description contradicted the code, it chose to stop and ask, while Neox and Codex decided on their own. The problem was our scoring, which rewarded going ahead and penalized asking first.
We excluded this batch from capability scores, added an unattended clause to the prompt and reran, and reported both results. We also noted that optimizing Neox against this scoring would push it toward "never ask", while in interactive use asking first is often right — Neox needs to support both behaviors.


