---
title: "From line numbers and hashes to a read ledger: seven changes to Neox's edit tool"
date: "2026-09-08"
tag: "Tools"
lang: "en"
reading_minutes: 18
source: "https://neox-dev.com/blog/edit-tool-from-hash-to-cache-coherence"
alternate: "https://neox-dev.com/md/blog/edit-tool-from-hash-to-cache-coherence.md?lang=zh"
---

# From line numbers and hashes to a read ledger: seven changes to Neox's edit tool

> Editing files is where agents fail most. Between March and September this year Neox's edit tool changed seven times: line numbers with expected_hash, then four layers of anchor fallbacks and a snapshot table, then a read ledger, content addressing, coherence triage, and finally the invariant "the ledger mirrors the context". This post lays out each version's design, problems, data and pseudocode in order, and explains why we concluded that editing is a cache-coherence problem, not a search problem.

When an agent writes code, reading files and searching code rarely go wrong. Editing is where it breaks: the model has to state exactly "replace this with that", and all it knows about the file is the copy it read in some earlier round.

Between March and September this year, Neox's edit tool changed seven times (see the figure above). Below, in order: how each version was designed, what went wrong, how we analyzed it, and what we changed.

## 1. Version one: line numbers + expected_hash (March)

### Why line numbers at first

Editing started as `old_string / new_string`: the model copies the text to replace. In March, wiring up GPT-5.x, we measured a problem: when editing large blocks, GPT echoed the whole old block verbatim — one edit took 97 seconds and 5,315 output tokens. Output is the most expensive and slowest part (decoding is serial), so the call at the time was: **have the model give line numbers instead of echoing old code**.

On March 27 the schema became this, with `old_string` hidden from the schema (still accepted in code, but invisible to the model):

```ts
edit_file({
  file_path,
  start_line, end_line,   // which lines
  new_string,             // replace with what
  anchor?,                // a snippet of the target line, to guard against drift
  expected_hash?,         // whole-file sha256 from the read, to avoid editing a stale version
})
```

Line numbers drift (earlier edits, linters and formatters change line counts), so we needed a way to tell whether the model's version was still current. We chose a whole-file hash: readfile returned content plus `expected_hash`, and edit sent it back for comparison.

### What one edit actually did

![Old design: one edit with line numbers + expected_hash](/site/blog/flow-edit-old-en.svg)

The design ran for months, and problems kept coming:

**1. Reading 5 lines meant two full reads and a sha256.** readfile only wanted to show 5 lines, but to attach a whole-file hash it read the entire file again to compute sha256. Every read was "2 full reads + 1 full hash", worse the bigger the file.

**2. Two edits in one file: the second always failed.** The model read once, got a hash, then sent two edits. After the first, the file changed and the second's hash was stale. On May 29 we checked `tool-trace.log`: all 14 `stale_snapshot` errors had `anchor=false, expected_hash=true`. DeepSeek never sent anchors and the system prompt told it "don't re-read files you've seen", so it reused one hash for several edits and every one after the first failed.

**3. Multi-line anchors always failed.** On May 25 a user reproduced it repeatedly in Java and Vue projects: the model passed several lines from readfile as the anchor, and the tool used a single-line `includes`, so a multi-line string never matched. We added four fallback levels:

```ts
function resolveAnchor(lines, startLine, anchor) {
  // L1 multi-line exact: split anchor by line, trim-compare from startLine   (~95% of cases)
  // L2 first line only: the model copied more or fewer following lines
  // L3 relocate within ±20 lines by first line: external edits shifted lines
  // L4 full-text search (requires expected_hash to pass): last resort against duplicate lines
}
```

**4. The snapshot table overflowed.** To fix problem 2 we added a snapshot table: before each edit, store "current content + its hash"; when the model arrived with an old hash, find that region in the old version and relocate it in the new file. In long sessions touching many files, the table (capacity 48, first in first out) evicted old snapshots the model was still reusing, producing "stale" errors though nobody had touched the file. Raising it to 256 with LRU only eased it.

### Our analysis

By the end of May the chain was "line numbers → whole-file hash → four anchor levels → snapshot table → offset tracking", each layer added to fix the one before.

Looking at them together, we saw they all did one thing: **use content to guess whether the model's copy was still usable.** The hash compared whole-file content, the anchor compared local content, snapshots compared past content. But "is the copy still usable" is a fact you can check directly: when the model read it, which version it read, and whether the file changed since.

## 2. The read ledger (July 15)

First we fixed the most direct performance problem — re-reading the whole file to compute a hash — by keeping records: what readfile read and which version, in a per-session ledger.

```ts
interface ReadEntry {
  rangeKey: string;     // 'FULL' or 'R:120-180'; each range of a file gets its own entry
  content: string;      // the text the model actually saw (not the whole file)
  startLine: number;    // absolute line number of content's first line
  lineCount: number;
  mtimeMs: number;      // fs.stat mtime at read time — the primary version stamp
  sizeBytes: number;    // byte size at read time — covers mtime's resolution gaps
  readAtTurn: number;
}

// sessionId → path → rangeKey → ReadEntry, at most 100 sessions kept (LRU)
const ledger = new Map<string, Map<string, Map<string, ReadEntry>>>();
```

Why `mtime + size` and not a content hash: we only need "did it change", not "what did it become". `fs.stat` doesn't read content and is an O(1) system call; a content hash requires reading the file first — reading in order to decide whether to read. Size covers mtime's resolution: within the same millisecond mtime may not change, but a different size definitely means a change.

With the ledger, readfile gained a short circuit:

```ts
function readfile(path, range) {
  const st = fs.statSync(path);                         // stat only, no read
  const hit = findFreshRead(path, range, st.mtimeMs, st.size);
  if (hit) return stub(`unchanged; the read from round ${hit.readAtTurn} is still valid`);
  const content = readRange(path, range);
  recordRead(path, { rangeKey: keyOf(range), content, mtimeMs: st.mtimeMs, sizeBytes: st.size, ... });
  return content;
}

// hit = same range + same mtime + same size: three O(1) comparisons
// On August 6 we added covering hits: a previous read that fully contains this range also hits (containment only, not overlap)
```

The same commit removed "re-read the whole file to hash": the hash was computed from the text already read, with zero extra I/O. For a 30-line file, a repeat read went from 1,281 characters to 159; for a 120-line file, from 2,284 tokens to 53.

The next day we closed a gap: `search` results carry line numbers and suggest `readfile(anchor_lines=[...])` for the hit blocks, but multi-block reads returned early before reaching the ledger — no dedup, no record — so a later edit couldn't fetch "the part the model saw". After the fix, each displayed block is recorded as `R:start-end`.

## 3. Content addressing (July 16)

With the ledger holding "the text the model actually saw", line numbers were no longer necessary. On July 16 we moved editing back to `old_string / new_string` and removed line numbers, the whole-file hash, anchors, snapshots and offset tracking — 856 lines deleted, 345 added.

```ts
function edit(path, old_string, new_string, replace_all = false) {
  const text = read(path);
  const hits = findAll(text, old_string);
  if (hits.length === 1 || (replace_all && hits.length > 0)) return applyAndWrite(...);
  if (hits.length > 1) return error('ambiguous_match', { lines: hits.map(lineOf) });
  return triage(path, old_string);   // no match → see next section
}
```

This is safe: if `old_string` matches uniquely, the model's snippet matches the disk byte for byte. Line drift doesn't affect it; no hash or snapshot needed.

The output-token concern was eased in two ways: the model only needs to copy a snippet long enough to be unique, not the whole block; and if the model gives only `start_line / end_line`, the tool cuts that range from the ledger and uses it as `old_string` (we call this the line-number bridge):

```ts
if (!old_string && start_line) {
  const read = findCoveringRead(path, start_line, end_line);  // full read first, else a covering range read
  if (!read) return error('need_old_string');
  old_string = sliceLines(read.content, start_line - read.startLine, end_line - read.startLine);
}
```

## 4. Coherence triage (July 24)

After content addressing, failure had one face: `old_string` not found. At first we kept three fuzzy levels (trailing whitespace, indentation normalization, similarity) — if not found, "find the closest".

On July 24 we re-analyzed the problem and concluded:

> The file content in the model's context is a cached copy of the file. If `old_string` matches, the copy is still valid; if it doesn't, the question isn't "how do we find it approximately" but "does the model's copy still match the disk?" That can be checked directly, without guessing from content.

A missing `old_string` has only three possible causes, each needing a different action from the model:

![An edit today: zero-cost fast path, triage only on failure](/site/blog/flow-edit-triage-en.svg)

```ts
function checkCoherence(path, mtimeMs, size): 'fresh' | 'stale' | 'unread' {
  const reads = ledgerOf(path);
  if (reads.length === 0) return 'unread';                     // never read this session
  if (reads.some(r => r.mtimeMs === mtimeMs && r.sizeBytes === size)) return 'fresh';
  return 'stale';                                              // read, but the version doesn't match
}

function triage(path, old_string) {
  const st = fs.statSync(path);
  switch (checkCoherence(path, st.mtimeMs, st.size)) {
    case 'unread': return error('file_not_read', 'read this file first');        // old_string was made up
    case 'stale':  return error('stale_read', 'file changed since you read it'); // linter / other process / last edit
    case 'fresh':  return error('string_not_found', nearestSnippetWithLineNumbers(path, old_string));
                   // read and unchanged ⇒ miscopied; show the real snippet with line numbers, say "no need to re-read the file"
  }
}
```

Here we deliberately differ from Claude Code. Claude Code checks **before every edit**: refuse if the file wasn't read, refuse if it changed after reading. Our reasoning: a unique exact `old_string` match already means the copy is consistent, so checking again at that point only costs efficiency — the model sees a line through search and edits it exactly, which is entirely correct. So we run the coherence check **after an exact match fails**, as triage: the fast path has zero added cost, and triage replaces the old fuzzy guessing.

We also removed the similarity level. Its reason to exist was "the file may have changed, the model may misremember" — exactly what triage now separates explicitly. Its cost was real: a legitimate rewrite (`const` to `let`, similarity 0.889) and a fabricated literal (`"xxx"` to `"bye"`, 0.846) differed by 0.043; any threshold was a bet, and losing it meant editing the wrong place. Only two unambiguous recoveries remain: whitespace-only differences (`blank_insensitive`) and `// ...` eliding a middle section (`elided`).

## 5. Refresh after writing, don't invalidate

The same day we found one part of the ledger pointed the wrong way: after a successful edit, we invalidated every read record for the file. So the most common path, "edit, then edit again", re-read the whole file every time. Yet the model knows the content after the edit — `new_string` is what it wrote.

```ts
function refreshReadsAfterWrite(path, newContent, previousContent, st) {
  const reads = ledgerOf(path);
  if (reads.length === 0) return;               // never read → add no read record
  if (reads.has('FULL')) {                      // read it all + knows what it changed ⇒ knows the current content
    reads.set('FULL', { content: newContent, mtimeMs: st.mtimeMs, sizeBytes: st.size, ... });
    return;
  }
  // Only partial reads: before August 13 all were dropped (re-read rather than over-claim);
  // since then, judged by the changed span — entirely before: kept; entirely after: shifted;
  // fully containing the change: re-sliced from new content; partially overlapping: dropped
  const span = diffLineSpan(previousContent, newContent);
  for (const r of rangeReads) keepShiftOrDrop(r, span);
}
```

Claude Code refreshes the whole file after writing, which it can do because it requires a full read before any edit. Neox has no such gate; refreshing to a full read when the model had only seen part of the file would make the next readfile collapse into a stub, hiding content it never saw. So we refresh only the ranges the model actually read.

## 6. The ledger must mirror the context

The day after triage shipped we found a hole in `fresh`: it checked whether the file had changed, not whether the model's copy was still in context. Compaction clears, truncates or folds readfile results into a summary; afterwards the ledger still said `fresh`, and triage told the model "you read it and it hasn't changed, you miscopied, no need to re-read" — when it could no longer see that content.

That gave us the ledger's invariant: **the ledger must mirror the model's context, not the disk.** If the ledger has a piece of content, the model must actually have it in front of it. Checking against this invariant, four places needed handling:

![Four faces of one invariant](/site/blog/flow-ledger-invariant-en.svg)

**Compaction (July 25).** Pair readfile results before and after compaction by `tool_call_id`; any change (deleted or truncated included) invalidates that file's ledger entries:

```ts
function findEvictedReadPaths(before: Message[], after: Message[]): string[] {
  const afterById = new Map(after.filter(isReadResult).map(m => [m.tool_call_id, m.content]));
  return before.filter(isReadResult)
    .filter(m => afterById.get(m.tool_call_id) !== m.content)   // vanished or rewritten = evicted
    .map(m => pathOf(m));
}
// All three compaction entry points (auto / manual / prompt-too-long recovery) call it, then invalidateReads
```

**Process restart (July 25).** The ledger lives in memory, while Neox is a long-running service with resumable sessions: reopening the app restores message history from the database, but the ledger is empty. The model's context clearly contains earlier reads, yet it would be told "not read, read it first". History doesn't store read-time mtimes, so the check uses content:

```ts
// current disk content appears verbatim in the output shown to the model ⇒ unchanged ⇒ current stat equals the read-time stamp
if (shownOutput(stripLineNumbers).includes(currentDiskContent)) recordRead(path, FULL, stat);
else keepUnread(path);   // changed → stay unread, which is the safe state
// full reads only; range / symbol reads skipped; files ≤ 512KB; only the last 40 reads
```

**Sub-agents (August 6).** We had assumed per-session ledgers isolated sub-agents automatically. Testing on August 6 showed sub-agents reused the parent session's ID: the parent read `store.js`, the sub-agent read the same file, hit the dedup and got "the earlier read is still valid" — but the sub-agent had a brand-new context with no such earlier read. One sub-agent hit this three times before working around it with search. The fix gives sub-agents their own read scope.

**Successful writes.** The "refresh after writing" from the previous section.

## 7. September: making failure cheap

On September 8 we tallied 501 edits: 426 succeeded, 75 failed. Of the 75, **70 were `file_not_read`** — 52 couldn't even find a "closest snippet" — and 5 were `string_not_found`. The bulk of failures wasn't miscopying but the model, eager to parallelize, inventing `old_string` for files it hadn't read. Each failure cost about two extra round trips.

Our analysis found two problems on our side:

- search results are line-numbered source, so editing from them is perfectly reasonable, but the ledger only counted readfile, so they were judged `unread` and denied any recovery. We now record lines shown by search as `S:start-end` range entries;
- refusals didn't include the source, forcing another read. Now, on a miss, lines containing `old_string`'s identifiers are shown verbatim with line numbers so the model can copy them next time.

Two more: whitespace-only recovery now also applies to `unread` / `stale` (that isn't guessing content); and an `old_string` of 40+ characters appearing identically at most 3 times is applied everywhere, with each location named in the receipt (systematic refactors need every occurrence changed; short strings still get the ambiguity error).

## Conclusion

After these seven changes, our conclusion is: **editing isn't a search problem; it's a cache-coherence problem.** The file content in the model's context is a cached copy, and an `old_string` match means the copy is valid; whether it is can be checked directly with the ledger and one `fs.stat`, with no need to guess via hashes, anchors, snapshots or similarity.

Based on that, Neox's edit path today is: reads go into the ledger (version stamp = mtime + size, no content read); edits match content exactly and write immediately on a hit, with no checks added; misses are triaged as `unread / stale / fresh` with three different responses; and the ledger follows the model's context through compaction, restarts, sub-agents and writes.

Compared with Claude Code, we borrowed its principles of refusing unread edits, refusing stale edits and refreshing after writes, but re-derived two of them for our architecture: the coherence check runs after a failure rather than before every edit, keeping the fast path free; and the ledger is rebuilt after process restarts, because Neox is a long-running service with resumable sessions.
