readfile had 24 parameters and models used 3: auditing tool parameter usage
Measuring tool parameter usage from real logs: readfile had 24 parameters, models used 3.
There's a natural impulse in tool design: the more complete the better. Reading a file by line range, by symbol, by regex, by function, in multiple ranges… every parameter comes with a reason.
On July 26 we decided to measure: which ones do models actually use?
The data
- 5,307 tool-call log entries, including 988 real tool calls across 876 inference rounds;
- 718 real requests sent to the model, with full tool definitions;
- 397 readfile calls;
- all on DeepSeek V4 Pro / Flash.
The limits up front: this is an observational audit. It can show "most parameters go unused", not "removing them makes things faster"; and it's a single-model sample, so stronger models need separate checks.
Finding 1: 24 parameters, 3 in use
| Parameter | Usage |
|---|---|
start_line | 73.8% |
file_path | 51.1% |
path | 48.9% |
num_lines | 40.1% |
end_line | 23.4% |
read_all | 8.8% |
anchor_lines | 4.5% |
pattern / symbol / ranges | 1.0% each |
| Everything else | 0.3% or 0 |
22 parameters were never used. By locating method, 92.5% of calls used the three most basic: a line range, the whole file, or just a path with defaults. For the other five methods we paid decision cost and tokens on 100% of calls.
Two parameters — num_line and change_context — weren't in the schema at all; the model invented them. A direct sign the schema was too complex.
Finding 2: aliases had the model flipping coins
file_path 51.1%, path 48.9%. Almost exactly half and half.
That's not a preference; the model doesn't know which to use and picks between two synonyms at random. The edit tool accepted four synonymous path parameters; search had four synonymous query names.
These aliases were each added for compatibility after a model used the wrong parameter name. But aliases make the schema muddier: four synonyms make the model less certain, wrong guesses rise, and another alias gets added. Fixing the description is what actually helps the model pick correctly.
Finding 3: the resident tool budget was badly inverted
Each request carried 20 resident tools, 29,832 characters of definitions — roughly 7,500–10,000 tokens.
| Tool | Share of definitions | Share of calls |
|---|---|---|
tool_search | 15.0% | 0.2% |
agent | 12.1% | 1.2% |
readfile | 9.8% | 40.2% |
open_surface | 9.8% | 0% |
bash_output | 8.6% | 0.1% |
search | 6.7% | 15.4% |
edit | 5.9% | 10.0% |
execute_shell | 5.1% | 24.9% |
- 6 tools were never called and took 19.9% of the budget;
- the three most expensive (
tool_search,open_surface,bash_output) took 33.4% of the budget for 0.3% of calls; - the four tools behind 90.5% of calls (read, shell, search, edit) took 27.5%.
The irony was tool_search: the entry point for dynamic tool loading, designed to save budget. Its description embedded the entire tool catalog, making it the single most expensive tool — used twice in 988 calls.
Finding 4: 1.13 calls per round — not the parameters' fault
| Calls per round | Share |
|---|---|
| 1 | 90.9% |
| 2 | 6.8% |
| 3 or more | 2.3% |
Mean 1.13: 988 calls over 876 rounds, nearly one to one.
At first we blamed decision load from complex parameters. Digging in, it was our own prompt. DeepSeek's dedicated supplement said "for coding tasks, proceed in order: identify relevant files → find the root cause → make a minimal change → verify → report", and "when parallel calls are supported, you may read in parallel". The generic version said "don't squeeze it out one call per round — if you need 5 files, send 5 reads in the same round" — which DeepSeek never received because it had its own supplement.
One was a sequential procedure with soft wording; the other was firm. DeepSeek got the former.
And we'd hit this exact trap two days earlier: a code comment cited "a model averaging 1.36 calls per round" as evidence, until we found the benchmark prompt at the time said "one write_file at a time" — the model was following instructions. On neutral tasks the same models send 20 tool calls in a round.
So 1.13 came mainly from our prompt, and that variable has to go before parameters can be judged. Also, GLM's supplement said "prefer reading with symbol" — we were actively pushing the model toward a mode used 1% of the time.
What we did about it
| Problem | Data | Action |
|---|---|---|
| Aliases | Proven harmful (51/49) | Remove |
| Never-called resident tools | Proven waste | Move out of the resident set |
tool_search inversion | Proven | Redesign or move out |
| Prompt suppressing parallelism | Contradicted by past measurements | Fix |
| Number of parameters and locating modes | Insufficient | Needs an A/B experiment |
Why the last one is insufficient: the log showed 0% failure for reads and edits, yet we knew edits really failed about 1.9% of the time. The "success" flag caught thrown exceptions, not returned error payloads. "Parameters caused no failures" can't be trusted — which is exactly why an experiment is needed.
Designing the experiment
- Experiment 0 first: remove the sequential procedure from the prompt and use the generic firm wording. A few lines of text. If total rounds halve, the decision cost of parameters will vanish into noise;
- Experiment 1: three tiers of read parameters — today's 24, a middle tier of 6, a minimal 3 (path, start line, line count);
- Unify aliases first, or all three tiers are contaminated;
- Run across strong and weak models — the answers will likely differ, and that difference is the final design: schema thickness adapted to model capability;
- The primary metric is total rounds and total time to finish; calls per round is secondary, and optimizing it directly repeats the "1.36" illusion; task success is a veto;
- The benchmark prompt must be neutral, with nothing hinting at a call rhythm;
- Environments must be isolated: two concurrent instances write the same log and the same database — that same day we'd contaminated an experiment this way. Run the headless CLI under a temporary HOME, not the desktop app.
Conclusion
The audit of real logs showed: 22 of readfile's 24 parameters went unused, aliases had the model picking randomly between synonyms, and a third of the resident-tool budget bought 0.3% of calls — the most expensive being tool_search, meant to save budget. The 1.13 calls per round came mainly from our DeepSeek prompt saying "proceed in order", not from too many parameters.
We dealt directly with the aliases, the never-called resident tools and the prompt suppressing parallel calls; the effect of parameter count itself lacked evidence and was left to a follow-up A/B experiment.


