---
title: "Should the agent resume on its own after a crash? We ran it for five months, then removed it"
date: "2026-09-23"
tag: "Engineering"
lang: "en"
reading_minutes: 7
source: "https://neox-dev.com/blog/crash-resume-built-then-removed"
alternate: "https://neox-dev.com/md/blog/crash-resume-built-then-removed.md?lang=zh"
---

# Should the agent resume on its own after a crash? We ran it for five months, then removed it

> When we rebuilt the core in April we made run state serializable, and in May used it for "if the process dies, resume automatically on restart". It sounded advanced. Months later it had produced a string of odd problems: the agent acting without the user saying anything, and an internal notice being taken as a user message and used as a session title. On September 23 we removed it: if it crashed, that turn is over — repair the history, mark it "interrupted", and let the user decide.

This is about a decision we reversed ourselves.

## April: state could be written to disk

When we rebuilt the agent core on April 18, we gathered the main loop's loose variables into a `RunState` and made it serializable: a snapshot after every round, and the agent could start from a snapshot.

In the notes we wrote then, crash recovery was the headline: "A 10-minute task, Electron crashes — before, redo everything; after, resume automatically." We even used the phrase "god-tier experience".

## May: resume automatically on restart

On May 12 we wrote the full design. The problem:

1. The user sends a message, the agent runs, messages are written to a local database as it goes;
2. The background service process exits for some reason (crash, out of memory, hot reload during development, manual restart);
3. The desktop app starts a new service process automatically, but the new process doesn't know a session was running;
4. The UI turns off streaming and the user has to resend. And by then the history has the user's message and a tool call **but no tool result and no complete reply** — a half-finished state — so after a resend the model sees a garbled context.

The design:

- A tiny new table recording only "which sessions are running, how far they got, which model". We didn't use April's `RunState`, because it only served the old execution path and the new runtime had no use for its fields;
- On service start, scan the table for interrupted sessions;
- **Repair the history**: if the last message is a tool call without a result, add an "interrupted, tool not run" result; if it's a half-written reply, drop it;
- **Automatically re-enter the conversation** so the model carries on, with "resumed automatically" shown in the UI.

The history repair was exactly right — the model is stateless and only sees the message history; if the history is valid, it works. The problem was the last step.

## A few months later

Automatic resume worked by inserting an internal notice into the session *as the user* after the service restarted (roughly "the service restarted, please continue") and starting a turn.

Over time these showed up:

- **The agent moved without the user saying anything.** As soon as the service restarted, the interrupted session started running on its own, editing files and running commands, possibly with nobody at the computer;
- **The internal notice was treated as a user message.** Automatic titling used it as material, and a session ended up titled "Resume service after restart".

Our first reaction was to fix the symptoms — for example, making titling filter out that internal notice. Even we felt it was wrong once it was done: we were adding filters behind a flawed path, and the next symptom would pop up somewhere else.

## September 23: crashed means over

The real question was: **should this recovery path exist at all?**

Once we asked that, the answer was simple. If the process died — crash, exit, power loss — **that turn is over.** Users want "it doesn't crash", not "after it crashes, something decides on my behalf to keep going". Automatic resume makes a decision for the user, and does it while they're not there.

The change that day:

- Keep **history repair**: close dangling tool calls, drop half-written replies, so the model sees valid history next time;
- Mark the session as interrupted, with one line in the timeline: "interrupted";
- **Never start a turn automatically.** It continues when the user sends the next message.

## Conclusion

A few months after automatic resume shipped, chasing a series of odd symptoms led us to conclude the problem wasn't any downstream feature but the recovery path itself: it carried on for the user while they weren't there and hadn't said anything, with the agent editing files and running commands. At first we only filtered the titling; later we confirmed that was just treating a symptom.

On September 23 we removed automatic resume, keeping only history repair and the "interrupted" marker, and the session continues when the user sends a message. History repair was solving a real problem, so it stayed.
