---
title: "A small feature took the agent 42 iterations: notes from our first month"
date: "2025-12-22"
tag: "Engineering"
lang: "en"
reading_minutes: 8
source: "https://neox-dev.com/blog/agent-42-iterations-first-month"
alternate: "https://neox-dev.com/md/blog/agent-42-iterations-first-month.md?lang=zh"
---

# A small feature took the agent 42 iterations: notes from our first month

> We rewrote Neox in Node on November 21, and by mid-December the desktop app and CLI both worked — but a small feature, "show the username top-right after login", took 42 iterations. Splitting the log by round, the waste sat in four places: fragmented exploration, failing edits, over-verification and extra files. This is how we fixed them one by one that week, including two bugs that took far too long to find.

We rewrote Neox in Node on November 21, 2025 (there was a Python version before). Within a month we had the CLI, the desktop app, adapters for several model providers, checkpoint rollback and a Plan mode. On paper it looked complete. In real use the worst part was speed: an ordinary frontend request took dozens of rounds.

These are notes from one week in mid-to-late December.

## A small feature, 42 iterations

The test task was simple: "Add user info to the frontend — show the username top-right after login, with a logout option."

It took 42 iterations. Splitting the log by round, the waste sat in four places:

| Rounds | Doing what | Wasted | Why |
|---|---|---|---|
| 1–5 | Learning the project | 3–4 | Looked at the tree, read the wrong directory, listed it, then read files |
| 16–28 | Editing | 12+ | `edit_file` failed to match, the model fell back to `sed`/`awk`, which broke too |
| 31–40 | Verifying | 8–10 | `grep` after every edit to check it landed |
| 41–42 | Wrapping up | 2 | Created `test_dashboard.html` and `DASHBOARD_UPDATE.md` that nobody asked for |

Fewer than half the rounds were actually writing code. That week our target was 15–20.

## Failing edits: the tool gave the model nothing to work with

The editing stretch wasted the most. When `edit_file` failed, it returned one line: "String to replace not found." From that the model cannot tell what went wrong — indentation, line endings, or code that had already changed. So it re-read the file, tried again, failed again, and finally reached for `sed`.

Two changes:

- **Looser matching, in a fixed order.** Exact match first, then ignoring trailing whitespace, then ignoring leading and trailing whitespace, then whitespace-normalized. Each level is looser than the last, and each still requires a unique hit. We also added a `change_context` anchor so the model can pass a short `old_string` plus "inside which function".
- **Return nearby content on failure.** On a miss, we search the file for the first line of `old_string` and return the candidate locations with a few numbered lines around each, plus "copy from here".

There was also a cosmetic problem: when adjacent spots in a file were edited several times, every diff redrew the whole block, so it looked like the model kept editing the same place. Whole-block replacement was swallowing the previous edit. `edit_file` now diffs by line and replaces only changed lines; if old and new are identical it skips the write and the diff entirely.

## A bug that took too long: the guardrail couldn't see the read

We have a guardrail: you must have read a file before you edit it. It stops the model from editing from memory.

For a while the model would read a file, then get blocked on the edit with "please read the file first". It obediently read again, edited again, got blocked again. Several rounds each time.

It took a day to find. The guardrail checked a read cache that the read tool wrote to on success. The read tool had been reimplemented and the cache write was never wired back in — the tool call succeeded, the call was recorded, but the cache the guardrail looked at was empty. On the evening of December 19 we finally found it and changed the guardrail to look directly at the last 20 tool calls for a successful read of that file (with normalized paths).

What we found: the guardrail read one store while the tool wrote another. It all looked fine until the tool's implementation changed, then it failed silently, and the symptom looked like "the model won't listen". We changed the guardrail to read the tool-call record directly.

## Another bug: loop detection never reached its top level

When the model called the same tool repeatedly, our loop detector had three levels: a gentle nudge on the 2nd call, a firm warning on the 3rd, a hard stop on the 4th.

In practice the message count went 29, 31, 33, 35, 37 and never stopped.

Two causes, both embarrassing:

1. **Detect before record.** When a loop was detected the code `break`-ed out, skipping the line that recorded the call. The detected call was never counted, so the count never reached 4. Now it records first, then detects.
2. **The nudge encouraged the loop.** For a `readfile` loop the message said "you can verify with readfile". The loop detector was feeding the loop. Now the advice depends on the tool: for a read loop, "the content is already in your context, use it"; for a search loop, try different keywords or read the file directly.

## Reading and searching: less paging

The second biggest waste was reading. `readfile` returned 200 lines at a time, so large files meant paging; search took one keyword and returned plain text, and the model had to pick line numbers and read again.

Between December 19 and 21:

- **Adaptive reads.** Files up to 700 lines come back whole; above that, a hint and the head of the file so the model locates first. New `list_matches` (all hit lines with one-line previews), `anchor_lines` and `ranges` (several fragments in one call, adjacent ones merged).
- **One name.** Reading had three names — `smart_read`, `read_file`, `readfile` — and the model and guardrails each recognized different ones. All became `readfile`.
- **One `search`.** Several keywords at once, ripgrep underneath, structured results with line numbers that feed straight into `readfile`, removing the "search, pick lines, read" round.
- **Exploration order.** The prompt now says: look at the tree once for the overview, then search by keyword recursively; don't list directories level by level.

Search also had a memory problem: searching a common word took the process from 47MB to 1.8GB. To show context around hits we read every matched file whole, then made a sanitized copy and a split copy — fifty large files is over a gigabyte. Now ripgrep returns context lines itself with `-C`.

## Requests too big, rate-limited

That week we also hit a rate limit of 50 requests per minute. One request in the log had 56 messages and 22.7K tokens, of which **21K were tool output**. Over ninety percent of each request was tool results.

That was the first time we saw that more rounds is not just slower — it multiplies cost and rate-limit risk: rounds times size per round. A lot of later work (truncating tool output, the read ledger, caching) started here.

## The prompt: from a paragraph to a decision tree

Finally, the prompt. The system prompt was 86 lines and said nothing about which tool to use when, what to do when an edit fails, or what counts as done. We added:

- **Tool choice.** For structure, look at the tree once; with a known path, read directly; search content with search; new files with write, existing files read then edit.
- **Editing.** Read before editing; copy `old_string` verbatim from what you read; after two failures rewrite the whole file; don't edit files with `sed`/`awk`.
- **Done.** Do what the user asked; don't create `.md` or `test.html` nobody asked for; a successful `edit_file` means it landed, no need to grep, verify at most once.

These look basic now, but each line matched rounds we could point to in a log.

## Conclusion

That week we split the 42-round log apart round by round and found that most of the circling wasn't the model's fault but came from tools giving too little information, or our own mechanisms getting in the way: a failed edit said only "not found", a guardrail read the wrong data, the loop detector's nudge told it to keep reading, and reads came 200 lines at a time.

In response we changed edit matching and failure output, fixed the guardrail and loop detector, let reads and search return more in one call, and wrote tool choice and completion criteria into the prompt. We have kept using this round-by-round log analysis since.
