We wrote 1,800 lines of "auto-test, auto-fix" and reverted all of it 90 minutes later
We wrote 1,782 lines of test-driven self-healing and reverted all of it 90 minutes later.
In an April capability review, the first gap we listed was: the agent says "done" as soon as it writes the code, and only starts fixing when the user comes back with "that's wrong". Output and verification were disconnected.
Drawing on the iterate loops of Devin and AmpCode and the SWE-bench evaluation style, we designed a "test-driven self-healing loop": every code delivery passes a "run tests → pass" gate, and the same machinery doubles as an evaluation framework.
1:59 pm: part one
The first commit added a new tool, run_and_iterate, with four components:
- Test command inference: recognizes vitest, jest, npm test, pytest, cargo, go and tsc; if none match, the command must be given explicitly;
- Failure parsing: pluggable parsers, with jest/vitest, pytest and a fallback that only looks at the exit code and the output tail;
- Error fingerprints: normalize the error and hash it; the same fingerprint twice in a row, or A-B-A-B alternation, means a loop;
- Runner: spawns the test command natively with a timeout cap and an output byte cap.
The model calls the tool once, gets a structured diagnosis and advice, edits the code, calls again with the same iteration ID, and repeats until it passes, a loop is detected, or attempts run out.
1,560 lines, 38 unit tests passing.
2:11 pm: part two
The second commit added a dedicated "self-healing sub-agent": when the main agent needs tests fixed, it sends a sub-agent with an isolated context to loop on the fix and hand back only a summary.
The sub-agent's tools were tightly restricted: read, edit files, call run_and_iterate; no deleting files, no git commit, and no shell — forcing it through our verification tool. Its prompt spelled out the procedure and hard rules: never bypass tests, never delete tests, never edit blindly. Temperature 0.2.
222 lines, 17 unit tests passing.
3:37 pm: everything reverted
Then we checked how Claude Code and Codex handle "run tests and fix bugs". The answer: neither has a dedicated tool. No test runner, no failure parser — just a plain shell, with raw output handed to the model.
Back in our own logs: the models we used already ran the tests after writing code, read the errors, fixed the code and ran them again. The model does this on its own. It doesn't need teaching, let alone a thousand-plus lines to "guarantee" it.
Once that was clear, the problems were obvious:
- It hard-coded an emergent model behavior as rules. Something the model already does got wrapped in a tool, a state machine and a dedicated role — an extra layer of translation between the model and the tests.
- Parsers can never be complete. We built in three test frameworks; there are hundreds. The fallback that "only reads the exit code" is just what a shell already is.
- Taking away the shell made it dumber. To force it through the verification tool we removed the shell, and then it couldn't even
lsa directory. - That's not where Devin's gains came from either. On a second look, their SWE-bench improvements came mainly from context management and planning, not from having a "self-healing tool".
Both commits were reverted in the same minute. We kept the design doc, with a note at the top: rejected — do not implement from this document.
Conclusion
Comparing Claude Code, Codex and our own logs, we found the model already runs tests, reads errors and fixes them after writing code on its own; the "test-driven self-healing" tooling hard-coded behavior the model already had, and taking away the shell made it dumber. So that afternoon we reverted both commits and marked the design doc "rejected".
For the "doesn't verify after writing" problem, we instead added half a page to the system prompt: "after changing code, run the relevant tests; it isn't done until they pass".


