Skip to content
← All posts

Same model, why does it feel better in Claude Code? Taking "feel" apart

Why the same model felt smoother in Claude Code, taken apart with real session data.

Data from one real session: 3.7 replies per turn, 82% empty bubbles, tool definitions taking 85% of each request

In mid-July we used Neox for an ordinary job: polishing the pages of a small project. The model was Grok 4.5, more than capable. But the experience came down to one phrase: it didn't feel good. Slow, fragmented, constant errors, and the page broke partway through.

Same model, same task in Claude Code: much smoother. This time we didn't say "the model's not good enough". We measured.

The data first

From a real session on our machine:

MetricMeasured
Messages the user sent22
Messages the agent sent82
Replies per user message3.7
Replies with no text at all67, i.e. 82%

Then the architectural problems dug up during the investigation, all reproducible:

  • Every request carried 133K tokens of tool definitions, 85% of the request. The shell sub-agent was carrying all 150 tools;
  • The context window was wrongly frozen at 128K while the model actually had 500K, so "you must compact" came up quickly;
  • One compaction freed only about 5K, because it could compress messages but not the 133K of tool definitions;
  • Parallel tool calls always triggered a "stream incomplete" error — a counting flaw in our parser;
  • "Retry" actually started a new session; the server had no "rerun this turn".

Put those numbers together and "slow, error-prone, unstable" is no illusion — it's a measurable property of the system.

What "feeling good" actually is

"Claude Code feels nice" isn't mysticism. We broke it into five observable dimensions:

DimensionWhat it isClaude CodeNeox then
Time to first progressFrom speaking to seeing real movementSeconds, straight to workLong — a ritual of reading, hashing, snapshotting first
Edit hit rateRight the first time, no reworkHighLow — line numbers and hashes go stale
Cutting lossesHow fast it changes course after a wrong turnFastSlow — easily stuck in protocol
Output densityUseful information per messageHighLow — 82% empty replies plus a running log
Self-repairCan it cleanly fix what it broke?StrongWeak — mechanical retries widened the damage

It feels good only when all five are high. Neox was being pushed down on every one.

Layer by layer

The editing protocol was the biggest source of friction, and the friction compounded. Neox's edits required: read the file for a hash → edit by start and end line → if the hash is stale or lines shifted, reread and retry. One edit changes the line count, every later edit's line numbers drift, so you change one spot, reread, change the next. N edits ≈ N rereads. Claude Code edits are located by content (anchored on a snippet of original text), so one change doesn't disturb another.

It was the same root cause as a parser bug we fixed that same week: position-based addressing everywhere — order, line numbers, counters — where the right answer is addressing by content or key.

Tools were too fine-grained, and the fragmentation spread to the output. Dedicated tools are controllable and auditable, but simple tasks got split into many small steps, and an empty reply could appear before and after each tool call. The timeline looked constantly busy, with very little information; what actually changed was buried among tool cards and empty boxes.

Every round paid tax on a mountain of tools. 133K tokens of tool definitions resent each round diluted the model's attention and pushed working context out of the window. Claude Code's tools are on call; it doesn't carry the whole mountain to every desk.

Failure recovery was mechanical retrying, not judgment. The incident chain in that task: a batch template edit broke the structure → a more complex script to patch it → a route conflict surfaced → closing tags fixed again. After the first failure it didn't fall back to "edit one file by hand"; it widened the damage. Verification checked HTTP 200 and nothing about the page structure. It retried the same direction repeatedly. Much of what makes Claude Code feel "cool" is cutting losses fast: two failures on one path and it switches. That's behavioral strategy, not model IQ.

Many errors were self-inflicted. Incomplete streams, retry becoming a new session, false "must compact", empty replies, half-applied edits — none of these were about the task being hard; the system made them. Each cost the user's trust and the model's tokens, adding up to "this thing always seems to be breaking".

One level deeper: two products optimizing for different things

Abstracting all of that:

What Neox optimized then   = traceability, recoverability, no overreach, no silent damage
What Claude Code optimizes = engineer feel, throughput, right first time, cutting losses

Each Neox "tax" was reasonable on its own: line numbers plus hashes prevent overwriting the user's changes; dedicated tools prevent dangerous commands; plans and verification evidence keep long tasks under control; giving every tool avoids crippling capability.

But our analysis showed that for a medium task like "polish a page", these costs outweighed the benefits. For large tasks, team work and audited settings the trade-offs are right; most daily work is small to medium, and users just want it done smoothly. Too many of Neox's actions served its own protocol, and users felt they were accommodating the tool.

What we fixed that week, and what's next

That week we fixed a batch of self-inflicted problems: the parser rebuilt around protocol indexes, the context window hot-updated to the real value, the shell sub-agent's tools cut from 138 to 37, 16 stray tools moved into on-demand loading, retry changed to rerun the turn reusing the original message, and empty replies no longer shown as empty bubbles (though the structural issue of creating a message for every step remained).

The three things that actually change the feel are next:

  1. Edits located by content, no longer depending on exact line numbers and whole-file hashes, with multiple changes submitted at once;
  2. Cutting losses: after two failures of the same kind, change strategy; before batch-editing more than 3 files, verify on one;
  3. Verify content, not just status codes: page tasks must check structure and key content, not just curl a 200.

After that: a fast path for small changes and the full process only for large ones; reports that state results — what changed, evidence, risks — with the process left in a collapsible timeline.

Conclusion

From real session data and a layer-by-layer investigation, we found that Neox felt bad mainly not because of the model but because of position-based edits, overly fine-grained tools, heavy tool definitions, mechanical retries after failures, and a batch of errors the system caused itself. That week we fixed the self-inflicted ones first; content-located edits, cutting losses and real content verification are next. When evaluating changes now, we also check whether a change adds steps the user has to do to accommodate the system.

View source