First lessons from supporting many models: every tool call needs its result
Four errors from supporting many models, all caused by tool calls and results not lining up.
Neox was multi-model from day one. In late November we connected the official Claude API and OpenAI, in early December Doubao and Gemini, and at the end of the month we adapted GPT to the Responses API.
The most common errors in that period were not "the model did badly" but requests rejected outright. Almost all of them came down to one thing: tool calls and tool results didn't line up.
1. We "tidied" thinking blocks into breaking
With interleaved thinking on, Claude would occasionally return:
thinking/redacted_thinking blocks ... cannot be modifiedand the session would die.
We had a message "normalization" step that merged two consecutive assistant messages into one — originally for compatibility with some endpoints — and a helper that stitched multiple thinking blocks together. Normally there are no consecutive assistant messages, but when a tool result was cleaned up or removed by pairing validation, two assistant messages ended up adjacent, got merged, the thinking blocks changed, and Anthropic refused.
Compared with Anthropic's rules and Claude Code's behavior:
| Anthropic requires | Claude Code | Us at the time | |
|---|---|---|---|
| Consecutive same-role messages | Allowed | Leaves them | Merged by hand |
| Thinking blocks | Must not change | Passed as-is | Merged several |
| Tool pairing | Strict | Drops orphans | Had it |
On December 21 we deleted the merging and the thinking stitching entirely and now send thinking blocks back untouched; the error hasn't come back since.
2. Parallel calls flagged as a loop, task killed
On December 26, GPT sent four parallel search calls in one round, with completely different patterns:
search pattern="app|main|flask|fastapi..."
search pattern="__name__\s*==\s*['\"]__main__['\"]"
search pattern="route\(|Flask\(|FastAPI\(..."
search pattern="api|endpoint|router|blueprint"The third got "you have called this 3 times with identical arguments", the fourth got TASK TERMINATED.
We planted this one. A week earlier, to fix "loop detection never reaches its top level", we had changed the order to "record, then detect". Fine for serial calls, but parallel calls in one round are processed one after another in a loop: the previous one was just recorded, so the next one's check counted it, and calls in the same batch interfered with each other.
The final approach: check every call in a parallel batch against the history before the batch, then record them together. The earlier "record was skipped" bug is fixed by always recording regardless of the check result, not by swapping the order.
This one came from changing the execution order in an earlier bug fix: two orders that give the same result serially don't in parallel.
3. Responses API: two IDs and a prefix
With GPT on the Responses API, the most common error was:
No tool output found for function call fc_call_5xI1...A Responses API tool call has two IDs: id (like fc_xxx, the output item's ID) and call_id (like call_xxx, the tool call's ID). When returning a result, function_call_output.call_id must equal the original call_id exactly.
We had an "ID normalization" step that added an fc_ prefix to anything not starting with fc_. So call_5xI1... became fc_call_5xI1... and the API could not find the call. Worse, the mapping lived in a global table that was never cleared between turns, so a repeated ID in a later turn was mapped to the old one.
Simple fix: never rewrite IDs; use exactly what the model sent. Normalization only generates an ID when there truly is none.
4. Unexecuted calls sent as history
On December 31 we hit the same error for a different reason. The log:
Sending function_call to API: {"call_id":"call_sgXM...","name":"delete_file"}
Found function_calls without outputs: ["call_sgXM...", ...]This time, tool calls the model had just issued — not yet finished — were sent to the API as history. Our flow saved the assistant message (with tool_calls) to history as soon as it arrived, and when converting to Responses format, marked every tool_call status: completed. If the user interjected or the tool run was interrupted, the next request carried "completed" calls with no outputs.
Codex puts a tool call and its output into history together and never sends a call without output. We now check during conversion and skip any function_call that has no matching output.
A related one: GPT's apply_patch arguments can be huge. Once we saw over 2,300 argument deltas stream in, more than 100KB in total, and never received the call's done event before the stream dropped. The tool never ran, and the next request failed with "no output found". Two additions: a size cap on arguments (warn at 100KB, stop at 200KB), and a check at stream end for calls that never finished — reported clearly instead of leaving a silent orphan in history.
Conclusion
Comparing each provider's requirements with our request logs over those ten days, we found the four errors had one cause: tool calls and their results didn't line up one to one.
We turned that into the first rule of our adapter layer: every tool call has exactly one matching result, with ID, content and order unchanged. In practice: don't merge or rewrite thinking, don't rewrite IDs, don't send calls without results, give interrupted calls an explicit "interrupted" result, and treat parallel calls as one batch.
Over the next six months we added more than a dozen model providers, and most format problems we hit traced back to one of these.


