The fastest-growing part of your context: tool results, and the ledger we built for it
Of the four context segments, system and tools are ones you can nail down. History is not — it only grows, and most of that growth is tool results. Re-reading the same 120-line file costs another 2,284 tokens without a ledger, or 53 with one. How we manage that segment: why the version stamp is mtime+size instead of a content hash, and one broken link we nearly missed.
The previous post covered the four segments of context, focusing on system and tools — the two you can nail down byte for byte.
History you can't. By definition it only grows, and most of that growth isn't conversation, it's tool results. Over one working turn, file reads, code searches, and command output all pile up here.
You can't freeze this segment. You can only keep duplicates out of it.
Reading the same file twice
The most common duplicate is a re-read. The agent reads pricing.ts to understand the tax logic, and a few turns later needs to edit it, so it reads again. The file never changed in between.
Without a ledger, that second read pushes the same bytes into context all over again:
On a real 120-line file we measured, the first read costs 2,284 tokens. On the second read, if mtime and size are both unchanged, the ledger short-circuits and returns a stub telling the model "unchanged, refer to the earlier result" — 53 tokens, 97.7% saved.
And it isn't just those 2,231 tokens. Once they're in history they get resent on every subsequent turn, and they push the context toward its compaction threshold.
Why the version stamp isn't a content hash
This step is easy to turn into a new cost of its own.
Our first implementation computed a sha256. Sounds rigorous; in practice it meant re-reading the entire file just to decide whether to re-read the file. The verification cost was the same order of magnitude as the thing it was saving — it cancelled itself out.
Now it's fs.stat for mtime + size, never touching the contents. A version stamp only has to distinguish "changed or not." It doesn't have to prove "what."
The hit condition is deliberately strict: same rangeKey, same mtime, same size — all three, or it's a miss. Better to miss and read again than to hit wrongly; serving stale content as fresh costs far more than one extra read.
The ledger is Map<sessionId, Map<path, Map<rangeKey, ReadEntry>>>, LRU-capped at 100 sessions. One file can hold several entries: reading lines 10–50 and later 300–400 are two different reads and must not evict each other.
A broken link we nearly missed
With this kind of ledger, what slips through isn't the main path — it's the branches that bypass it.
Our search already returns line numbers with hits and nudges the model to follow up with readfile(anchor_lines=[...]). The chain is well designed: search gives line numbers → read those blocks → edit.
Except multi-block reads (anchor_lines / ranges) hit an early-return before reaching the ledger. So that path neither deduped nor recorded anything. The damage wasn't just lost tokens — edit's line-number bridge pulls "the exact slice the model saw" out of the ledger, and with nothing recorded, the chain broke at the final step.
The fix was to route the multi-block branch through fs.stat too, recording one entry per displayed block keyed as R:start-end. Re-reading the same anchors now hits stubs across the board, and edit can pull content from the ledger. After a write, invalidateReads clears every block entry for that file.
Relatedly, this is why edit moved to content addressing: give it an old_string and let it match, rather than giving it line numbers. Line numbers drift; content doesn't. The old implementation's pile of hashes, anchors, snapshots, and line offsets is gone — a unique old_string match is the safety net, and rollback belongs to the per-turn ShadowGit commit, not to a second mechanism inside the tool.
Three things we deliberately didn't do
No content comparison for read-free verification. The comparison requires a read. Deadlock.
No cross-session persistent cache. The ledger dies with the session. Reusing it across days sounds great until you consider that stale data reuses just as well — another process edits the file between sessions and you'd never know.
No pushing for full parameter coverage. From the tool audit in the previous post: readfile ships 24 parameters and 22 were never used by the model; 92.5% of all locating is done with line ranges, read_all, or the default. The ledger serves those hot paths first — rare parameters aren't worth complicating the hit check for.
The two disciplines only work together
Prompt Cache keeps the bytes of that big leading prefix from jittering. ReadLedger keeps duplicates out of history. They're neighbors, not the same medicine.
Do only the first and you'll see a decent prefix hit rate while five full copies of the same file sit in your history and blow the context anyway. Do only the second and history stays clean while one timestamp in the system block makes the whole prefix recompute every turn.
Long sessions need both.
Worth separating from a nearby idea: this has nothing to do with "stuffing files into the prompt." RAG and vector stores answer "how do I find the relevant content." A ledger answers "stop hauling the same content around after you've found it."


