---
title: "GPT kept stopping halfway: we added a pile of gates, then removed them within a day"
date: "2026-03-04"
tag: "Engineering"
lang: "en"
reading_minutes: 8
source: "https://neox-dev.com/blog/gpt-stops-mid-task-completion-gate"
alternate: "https://neox-dev.com/md/blog/gpt-stops-mid-task-completion-gate.md?lang=zh"
---

# GPT kept stopping halfway: we added a pile of gates, then removed them within a day

> In early March, wiring up GPT-5.x, tasks that ran fine on Claude would often end right after GPT said "let me look at this file first". Our first instinct was a completion gate — no text-only endings, at least one tool call, bounce anything matching "let me first". Then even "hello" looped forever. After comparing with Codex we removed every gate by noon the next day; what actually helped was merging five system messages into one. Along the way we caught a "ghost retry" still running commands after a task had ended.

At the end of February we added GPT-5.3, and on March 1 built per-model-family configuration (Profiles) that parameterizes differences in prompting, completion checks and transport instead of scattering if/else through the code.

A problem showed up right away: **the same task that Claude finished end to end, GPT often abandoned halfway.** Typically it replied "let me look at the xxx file" or "Let me check the config" and the turn ended, with no tool call at all.

## First instinct: add gates

We already had a "completion gate": when the model wanted to finish, check for evidence of execution; if none, bounce it. GPT stopped more, so tighten the gate. On the morning of March 2 we wrote a plan expanding the completion part of the Profile from 4 fields to over 12:

- **Treat unclassifiable intent as an execution task.** Previously unknown intent passed as conversation; for GPT, make it strict — no ending without a tool call.
- **Grace for text-only endings.** If GPT replies with text only, refuse the first 1–2 times and inject "please use tools" each time.
- **At least one successful tool call** before finishing.
- **A batch of "intermediate phrase" regexes**: `Let me check...`, `I'll first...`, the Chinese equivalents — any match means "not done yet".
- **Stream anomaly checks**: replies that are too short, or `finish_reason=stop` with no tool call, treated as suspicious truncation and retried.

Each item made sense on its own, and part of it had already shipped the night before.

## Result: even "hello" looped forever

What broke first wasn't long tasks but small talk. A user typed "what can you do", "hello", any greeting; GPT replied with text — bounced by the gate, "please use tools" — GPT replied with text again — bounced again. It could never finish.

Plain cause: strict mode treated every input with unclassifiable intent as an execution task requiring tool evidence. Small talk doesn't need tools. We tried widening a regex whitelist for "this is chat" and quickly saw the dead end: **there is no end to the ways people chat.**

## Comparing with Codex

So we looked at how Codex handles it. Its turn loop is surprisingly simple:

- tool calls → run them, continue;
- no tool calls → the turn ends.

No completion gate, no text-only grace, no evidence checks, no nudging. It trusts the model to decide when to use tools and when to answer.

Around noon on March 2 we pushed two changes back to back:

1. Every GPT profile's intent fallback went back to lenient: only tasks explicitly recognized as execution, modification or analysis require tool evidence; everything else passes. The chat regexes from the night before were reverted.
2. The completion gate and text-only grace were deleted entirely, matching Codex. That commit removed 71 lines and added 13.

We kept one thing: **tool-call leak detection.** GPT occasionally emits a tool call as text (for example a garbled string containing `to=functions`), which our JavaScript stream parser would show straight to the user. That's a format problem, not a behavior problem. We kept the check and made it stop emitting text the moment a leak is detected during streaming, instead of cleaning up after the stream ends.

## So why did GPT stop halfway?

With the gates gone, the stopping problem itself remained. That afternoon we changed direction: instead of blocking at the exit, look at the input.

At the time our system prompt was 3–5 separate system messages: base role, tool rules, project notes, dynamic reminders… Claude didn't care about that structure, but Codex and OpenCode both give GPT **one single** system prompt.

The afternoon's changes:

- merge the system messages into one;
- add "keep going on your own" guidance modeled on Codex's prompt: don't stop to report before the task is done, don't ask "shall I continue";
- enable parallel tool calls for all GPT models;
- judge "wants to continue" by structure instead of keywords: a short reply, with tool calls, without a question mark, counts as intermediate.

Here we found that models differ a lot in how sensitive they are to prompt structure: the same multi-part system prompt was fine for Claude and made GPT quit early.

## A ghost retry caught along the way

On March 4, chasing something else, we found a scary timeline in the log:

```
10:54:44  task complete, isRunning=false
10:55:06  an earlier connection times out
10:55:07  retry: a new model request is sent
10:55:21  the new stream starts running tools
10:55:35  five execute_shell (curl) calls in a row
```

The task had ended, nothing showed in the UI, yet requests were being sent and commands run in the background, and ESC did nothing.

Root cause: when a task ended we set the abort controller to null **without calling `abort()` first.** The abort signal handed downstream never fired, so in-flight HTTP requests, retry sleeps and timeout streams all stayed alive; when that timeout hit, the retry loop launched a brand-new request.

The fix has three layers: call `abort()` before nulling the controller at task end; check the signal at the top of every iteration of the model request retry loop; check again before each retry.

## Conclusion

Comparing Codex's turn loop with our own logs, we found GPT stopped halfway not because of the finish check but because of the system prompt's structure: multiple system messages affected GPT a lot and Claude hardly at all. The completion gate we added first judged intent with regexes; small talk can't be enumerated, so even "hello" looped forever, and we removed it all the next day.

The final fix aligned with Codex: no completion gate, one merged system prompt, guidance to keep going, parallel tool calls on, and only tool-call leak detection kept. Along the way we also fixed a ghost retry caused by not actually cancelling requests when a task ended.
