Rebuilding the agent core: one pipeline, one state machine, concurrency decided per call
On April 18 we rebuilt the execution paths, the main loop and concurrency decisions in one day.
By April, Neox had a desktop app, a CLI, sub-agents and a collaborative mode, and features kept piling up. Internally the feeling was one word: "messy". A sharper version: "too weak — it relies on the model to check itself."
On April 18 we stopped and gave the core a checkup, listing exactly where the mess was.
The checkup: three sets of execution logic
The biggest problem: execution paths didn't converge.
- The main path the desktop app used (StreamedRunner) hand-assembled its own sequence: validate arguments, loop detection, mode filtering, a pre-execution reasoning check, permission checks, guardrails, parallel execution, write back results. No risk evaluation, and no hook after failures.
- The path the CLI and sub-agents used (agentLoop) assembled another: repair tool names, parse arguments, risk evaluation (reject high-risk outright), loop detection, call the tool directly. No permission check, and no guardrails.
- The repo also had a fully designed 6-stage tool pipeline. Neither main path called it. An island.
Five specific places:
- Loop detection lived in three places;
- Permission checks ran twice on the same path;
- Risk evaluation existed only on the CLI path; the desktop app had no up-front block for high-risk commands and relied on a guardrail blocklist;
- Three objects (error-pattern memory, tool-usage advisor, task-intent tracker) were constructed but almost never written to — shadow systems;
- The failure path had no post-hook, so telemetry, retry and fallback couldn't respond to failures consistently.
In other words, the same safety check existed on some paths and not others, while we had assumed every path had it.
The main loop: 15 loose variables
The CLI path's main loop was a while(true) maintaining fifteen variables at once: messages, budget, loop detector, force-terminate flag, loop-guide count, whether self-reflection had fired, error-pattern memory, recent text fingerprints, the round of the last loop check, model retry count, whether compaction had been tried, whether it had resumed, pause control, abort signal…
Whether to continue was decided entirely by if / break / continue and flags. A newcomer reading for five minutes couldn't say what stage the run was in.
The desktop path was somewhat better: each continue wrote a transition.reason, but that was a label applied afterwards, not a decision made up front, and each transition had to spell out all 9 fields by hand — two places handled the same field differently.
Concurrency: two allowlists, one blunt cut
Which tools could run in parallel was decided by global allowlists. Problems:
- There were two lists (19 entries and 13), plus a third place using different semantics for "read-only", all out of sync;
- The decision ignored arguments: reading a 50MB file and a 20-line file were both "safe";
git statusandrm -rfwere bothexecute_shell, both "unsafe"; - The batching algorithm: after the first unsafe tool, everything else runs serially. Five reads following it never got to run in parallel;
- Shared state (error memory, loop-detection history) was pushed to without protection during parallel runs, and there was no concurrency cap — dozens of parallel calls at once could spike memory.
April 18: twelve steps
We didn't rewrite in one go. We cut it into a dozen-plus commits, each runnable and revertible, done in a day:
Stop the bleeding (P0). One allowlist; a concurrency cap; defensive copies of shared state before parallel runs. No architectural change, just preventing damage.
The three essentials (P1).
- The abort signal threaded end to end. From the task down to every tool and every network request — cancel means cancel.
- Concurrency decided by arguments. Each tool implements
isConcurrencySafe(args): small reads can run concurrently, large reads drop to exclusive; read-only commands likels,git status,git diffandgit logcan run alongside other read-only tools. The allowlist becomes a fallback. - Interleaved batching. Consecutive safe calls form a parallel batch, an unsafe call gets a batch of its own, and safe calls after it start a new batch — no more cutting everything off at the first unsafe call.
Take the main loop apart (P1-2).
- First define three types —
RunState,Transition,Terminal— and move all fifteen loose variables intostate.xxx; - Then pull each recovery branch out into its own transition function: hard-loop guidance, delegation summary, runtime compaction, self-reflection, collaboration checkpoint;
- About 290 lines of model-call try/catch in the loop became
invokeLLMStream, down to 5 lines; about 300 lines of tool-batch handling becamerunToolBatch, down to 1 line; - Finally make
RunStateserializable: a snapshot after every round, and the agent can start from a snapshot.
Afterwards the main loop reads as "call the model → pick a transition based on the result → run it", and the current stage is obvious at a glance.
Conclusion
The checkup showed the core was messy because every new execution path got its own copy of the checks, and over time each path ended up missing something — the desktop path had no risk evaluation, the CLI path no permission check; the main loop ran on 15 loose variables; concurrency depended on two out-of-sync allowlists.
On April 18 we finished the refactor in a dozen-plus commits: one pipeline for all execution paths, the main loop rebuilt as RunState plus transition functions, and concurrency decided per tool from its arguments with interleaved batching. Everyday conversations barely changed, but a new recovery strategy is now one transition function, every entry point's tool calls go through the same checks, and run state can be written to disk.


