---
title: "Notes on what we optimized in the Neox agent this round"
date: "2026-09-30"
tag: "Engineering"
lang: "en"
reading_minutes: 9
source: "https://neox-dev.com/blog/agent-optimization-notes-rounds"
alternate: "https://neox-dev.com/md/blog/agent-optimization-notes-rounds.md?lang=zh"
---

# Notes on what we optimized in the Neox agent this round

> Neox always felt a bit less smooth than Claude Code, and we could not say why. This time we skipped the gut feeling: we rebuilt every round of the agent loop from logs and compared against Claude Code on the same model. Accuracy was fine. Rounds were not: the same jobs took 3 to 6 times as many. And the cause was not a weak model, it was the model copying the "one small step per round" rhythm in its own history.

Code written with Neox usually comes out right, but we had long felt it was slower and chattier than Claude Code, without being able to say exactly why. This time we decided to get the data first and then decide what to change.

## How we measured

Two things.

**First, rebuild every round of the agent loop.** Run real tasks through the CLI with debug logs on, then take each model call apart: which tool it called, with what arguments, what came back, how many tokens it cost, how much hit the cache. A task is ten or twenty rounds; laid out side by side, it is obvious which rounds did work and which were wasted.

No toy problems. We built a small ledger CLI (about ten files) with a few planted bugs and wrote longer tasks against it: reproduce and fix bugs from a user report with no failing test, switch to multiple accounts and migrate old data, add a web service with a page, make it bilingual, plus a read-only investigation in the Neox repo itself.

**Second, compare against Claude Code on the same model.** Both sides on Opus 5.5, same tasks, comparing rounds, time and test results. With the model held fixed, any difference is the agent.

## Result: accuracy is fine, rounds are 3 to 6 times higher

Tests passed on both sides, and the investigation answers mostly agreed. The gap was all in rounds:

| Task | Claude Code | Neox (before) |
|---|---|---|
| Fix three bugs | 4–5 rounds | 18 |
| Multiple accounts | 14–15 | 37 |
| Web service | 11 | 36 |
| Bilingual | 9 | 14 |

Reading the rounds made it obvious. Claude Code reads the whole project in round one with a single `git ls-files && cat ...`, then writes all the tests with one heredoc and runs them in the same command. Neox listed the directory, read files over three rounds, then changed one file per round, then spent another round running tests. **Under Neox, Opus sent exactly one tool call in each of 33 replies**; the same model in Claude Code writes three files in one round.

## Root cause: the model copies its history

We first blamed the prompt, and there was a real problem there: our tool rules said "read with readfile, not cat; keep the shell for builds and tests" and "independent reads may run in parallel" — which amounts to telling the model to change files one at a time, one tool per job. Fixing that helped, but unevenly: the same task took 11 rounds one time and 36 the next.

So we replayed Neox's real requests and changed one variable at a time:

| Changed | Result |
|---|---|
| Channel | An obviously parallel request on the same channel got 4 calls per round every time. Not the channel |
| Thinking | Thinking off or at maximum: still 1. Not it |
| System prompt | Removed entirely: still 1. Not it |
| Tool descriptions | Cut to one line: still 1. Not it |
| **Rounds in history** | Merging earlier rounds into "several calls per round" **immediately gave 2–3** |
| Parallel instruction | Claude Code's exact wording in the system prompt: barely any effect |

The model copies its own history. The first steps of a task are naturally one small action per round (list the directory, read a file, read another), the model locks into that rhythm, and no prompt pulls it back. That also explains the swings on the same task: it depends on whether the first round that writes files happens to go parallel.

Claude Code often sends one call per round too, but each call is big: one Bash reads every file, one heredoc writes several files and runs the tests. Fewer rounds follow from that.

## What we changed

Since prompting couldn't pull the model back to several calls per round, we took a different route: make each call able to do more. Specifically:

- **readfile accepts directories and globs.** `paths=["."]` reads a small project in one call, following git's file list and skipping binary, lock and very large files; the old cap of 8 files per call is replaced by a token budget.
- **Small projects get the file list up front.** That saves the "let me look at the directory" round. The list is taken once per session and frozen, so it never breaks the prompt cache.
- **write_file writes several files, edit changes several files, in one call.** Each file still goes through the full single-file path (snapshots, read ledger, empty-content guard), one failure does not stop the others, and the desktop shows a change card per file.
- **Files written by shell commands get change cards too.** Write-type commands look at the workspace before and after and list the changed files in the result. New files can now be written with a heredoc and `&&`-chained to the tests, like Claude Code does, without losing anything in the UI.
- **Resident tools cut from 23 to 17.** git status, git diff, listing directories and running tests — anything one shell command does — now load on demand. The finer the tools, the more the model does one thing per tool.
- **Verification: one end-to-end check after the tests pass.** The rule used to say "any failed step means not done", so the model treated a favicon 404 or a styling nit as a failure and kept fixing and restarting the server. After classifying every round, "still checking after green" dropped from 31 rounds to 6.
- **A nudge right before writing.** Prompts alone could not pull the model back, but one line at the moment it is about to start writing — "changes that do not depend on each other go in one call" — turned 3 of 4 replays into multi-file writes (0 of 4 without it). Timing matters: once it has already written the files one by one, the nudge is useless.

## Results

Same Opus 5.5 again:

| Task | Claude Code | Neox before | Neox after |
|---|---|---|---|
| Fix three bugs | 4–5 | 18 | 11–15 |
| Multiple accounts | 14–15 | 37 | 16 |
| Web service | 11 | 36 | 18–23 |
| Bilingual | 9 | 14 | 7–11 |

Multiple accounts and bilingual are level now; the other two went from 3–6× to about 1.5×.

Wall time is still about twice as long, but now we can say why: our side ran through a relay that streams just over 30 tokens a second against 100-plus on the official channel, with roughly 4 extra seconds before the first token on every request. On the same task Neox produced fewer tokens than Claude Code and still took twice as long, so the time gap came mainly from channel speed rather than the agent itself.

When the Opus budget ran out we switched to DeepSeek V4.1 Flash and compared Neox with itself, two runs per task to see the variance:

| Task | Before | After |
|---|---|---|
| Fix three bugs | 11 | 8 / 8 |
| Multiple accounts | 22–27 | 9 / 11 |
| Web service | 16–18 | 12 / 14 |
| Bilingual | 15 | 8 / 7 |

Variance also dropped a lot: the same task used to take 7 rounds one time and 21 the next; now two runs differ by 0 to 2.

## Watching the cache along the way

We tracked cache hit rate on every run: 86–95% overall, and the new nudges and file list are appended at the end, never touching the prefix. Two things came up:

- **The relay dashboard's fixed 20K-plus "uncached input" per request is not real.** That is the size of our system prompt plus tool definitions. The usage returned upstream at the end shows it all served from cache; the dashboard just displays the estimate from the start of the request. The prefix is on the large side, which is part of why six resident tools went.
- **Unlocking a deferred tool mid-session threw away a whole cached prefix.** The tool list sits at the very front of the prompt, so adding one tool invalidates everything after it — about 30K tokens each time in our runs. Now a newly unlocked tool is called through `call_tool` with the tool list unchanged, and joins the list only when the cache would be rebuilt anyway (the next turn after five or more minutes idle).

## Conclusion

By replaying the logs round by round and comparing against Claude Code on the same model, we found the gap between Neox and Claude Code was not accuracy but rounds, and that the extra rounds came from the model copying the "one small step per round" rhythm in its own history — not from the channel, the thinking level or the prompt.

Parallel instructions in the prompt and zsh notes in the tool descriptions made little difference. What worked were changes to the mechanics: directory reads, multi-file writes and edits, a file list up front, a nudge at the moment writing starts, and aligning zsh with bash in the executor.

After the changes, the round gap on the same Opus went from 3–6x down to 1–1.5x, and on DeepSeek V4.1 Flash rounds dropped and variance shrank. We also found that the variance of a single run was larger than the effect of the changes, so every result here was run at least twice. The remaining gap is mostly in bug hunting and the web-service task, which is what we look at next.
