---
title: "ReAct is a product of 2022: re-deriving the agent loop from model physics"
date: "2026-07-04"
tag: "Engineering"
lang: "en"
reading_minutes: 10
source: "https://neox-dev.com/blog/react-2-attention-compounding"
alternate: "https://neox-dev.com/md/blog/react-2-attention-compounding.md?lang=zh"
---

# ReAct is a product of 2022: re-deriving the agent loop from model physics

> In early July we stopped to ask: ReAct's "think → act → observe" loop was designed for 2022 models with 4K context, no thinking channel, and only one reliable tool call at a time. All three changed by 2026 — so how should the loop change? This is the record of that derivation: from eight physical facts about models to "attention, not tokens, carries the compounding", plus a first draft we rejected ourselves.

After most of a year building agents, we increasingly felt that many design choices were made out of inertia. In early July we stopped and worked it out from scratch: **how should ReAct actually be "new"?**

This is a theory draft, not a shipped feature. We're publishing it because it shaped many concrete decisions afterwards.

## ReAct's three designs were 2022 patches

| Original design | Reason in 2022 | Reality in 2026 |
|---|---|---|
| Reasoning as `Thought:` text | No thinking channel; reasoning had to live in output | Native thinking channel — measurable, adjustable, separate from the answer |
| One action per step | Models could reliably produce only one action | Native parallel tool calls, several per round, long reliable |
| Observations appended verbatim | 4K context; if it didn't fit, too bad | 100K–1M context; the problem moved from "doesn't fit" to "attention diluted" |

Conclusion: ReAct's **iterative semantics** are right — try, look, correct; context accumulates and hit rate rises round by round. That must stay. But all three **mechanical implementations** are obsolete. What needs upgrading isn't the loop but how each step inside it is implemented.

## Eight physical facts about models

We started only from verifiable engineering facts:

1. **Reading is cheap, writing is expensive.** Prefill is parallel (thousands of tokens per second), decoding is serial (tens per second). Reading 1,000 tokens costs about as much latency as writing 10–20. So loop time is dominated by **output**.
2. **Cache economics.** An unchanged prefix is nearly free to reread; change the middle and everything after is recomputed. Context should be append-only.
3. **Attention dilutes.** Under softmax normalization each of N tokens gets a baseline of about 1/N; relevant signal must beat noise to be used. "Context rot" isn't about length — it's **signal-to-noise falling below a threshold.**
4. **Position has geography.** The prefix is stable and cacheable; the end has the most influence. Where information sits matters as much as what it is.
5. **Models are stateless.** Every round rereads everything. The "agent" is really the context; building an agent means managing the context's lifecycle.
6. **One sample isn't the ceiling.** Verifying is cheaper than generating. High-stakes decisions merit several samples and a pick.
7. **Thinking compute is adjustable.** Thinking tokens buy decision quality, with different marginal returns per step.
8. **Context is a training set.** Examples in context are gradient-free training data — that's why hit rate rises over rounds. Conversely, **models imitate their own past mistakes in context**; raw failure transcripts left in place invite repeats.

## The core conclusion

From these eight facts we concluded that the per-round gain in "getting the next step right" depends on attention signal-to-noise (decision-relevant tokens / all attended tokens), not to context length.

- Naive ReAct: observations pile up raw, noise grows linearly, SNR falls, compounding decays, and it rots;
- The better approach: curate every round so signal grows faster than noise, keeping compounding steady or rising.

So what we need to design is no longer "how to order the steps" but "what the model sees at the moment it decides".

## A new five-stage loop

From this we split each round into five stages, which we call Curate-Act:

- **Shape.** Tool results don't go in raw: on entry, extract what matters for the decision and keep a pointer to the original, expandable at any time.
- **Place.** Three zones: cold is the prefix (system prompt, task charter, stable knowledge — unchanging, fully cached); warm is the body (an append-only log — the "training set"); hot is the tail (a small working set regenerated each round: current goal, active constraints, latest facts, open items, under 2K tokens).
- **Decide.** Thinking budget follows difficulty: think little on routine steps, more on errors, failed verification and forks in the plan, sampling several times and verifying when needed.
- **Act-wide.** Every decode has fixed overhead, so emit all independent actions in one round. Move whatever can go into input (plans, templates, state) into input; output only decisions and calls; ask for diffs rather than full text.
- **Metabolize.** Mark superseded facts stale; rewrite failures as one-line lessons — "tried X, failed because Y" — instead of leaving the raw failure on stage; clear duplicates in a later batch compaction.

On curating "raw material" we later added a limit: raw material kept side by side has value of its own. Attention is content-addressed; a later fact may connect with earlier material that "looked irrelevant at the time", and no summarizer can predict that. So the model's own reasoning, choices and lessons are never compressed; only bulky observations are, and only once the window gets tight, oldest first.

## The first draft we rejected

Section 9 of the first draft proposed a "layered reflex arc": like the brain, let small models or rules triage first, so routine matters never reach the main model and only "surprises" do. The main model would be involved less and less, and cost less and less.

After analysis we rejected it, for two reasons:

- **Its gain was fewer tokens, not more capability.** All the gains of layered routing are cost. We wanted maximum capability; cost belongs in the user's budget controls, not in the cognitive architecture.
- **The routing paradox.** "Is this worth the main model?" is itself one of the hardest judgments. The gatekeeper must be as smart as what it guards, or it becomes the system's capability ceiling. The most dangerous failure is a big thing that "looks routine" being silently swallowed by a weak router — with no audit trail.

## Three design principles

After rejecting it we redefined the architecture's role: it doesn't arrange each step for the model; it defines which tools the model can use, what context it sees and how reality responds (deterministic checks such as tests and compilers), and leaves the rest to the model's judgment. Each new mechanism is checked on whether it helps the model or decides for it. Concretely:

1. **All judgments go to the strongest model.** What to do, when, whether it's done, whether to think harder — the main model decides. External rules may escalate (e.g. force more thinking after an error) but never de-escalate or block.
2. **Helper components do three kinds of work and fail open.** Execute the main model's spec (correctness judged by tests and compilers); gather more evidence (never decide "no need to look"); compress context (reversibly — the main model can always expand the original). We don't build components that can make the main model silently miss information.
3. **Capability comes from four factors.** How much evidence each decision gets, iteration speed, signal-to-noise, and sampling several times at key points. Saving tokens is a by-product of these, not the goal.

## Falsifiable predictions

A theory only means something if it can be proven wrong. We listed five experiments:

1. Same task, same model: raw accumulated context vs curated context — hit rate over rounds;
2. Same token budget: inject "relevant summaries" vs "raw output" — next-step accuracy;
3. Full-text output vs diffs — is quality equal and time lower?
4. High thinking throughout vs difficulty-based levels;
5. Failures kept raw vs rewritten as lessons — repeat-mistake rate.

If any fails, the corresponding law goes back to the drawing board.

## Conclusion

Re-deriving the agent loop from the physical properties of models, we concluded that the loop's iterative approach should stay, but each step should be redesigned around the signal-to-noise of what the model sees when it decides; all judgments go to the strongest model, and mechanisms only supply tools, context and verification.

By this conclusion, the GPT completion gate removed in March and the self-healing tool reverted in April were both mechanisms deciding for the model. This is a theory draft; the five experiments above haven't been run yet, and the conclusions will be revised by their results.
