---
title: "Running the agent on desktop for 10 hours: first we fixed 6 bugs in our own monitor, then 3 P0s"
date: "2026-08-01"
tag: "Engineering"
lang: "en"
reading_minutes: 10
source: "https://neox-dev.com/blog/overnight-soak-desktop-agent"
alternate: "https://neox-dev.com/md/blog/overnight-soak-desktop-agent.md?lang=zh"
---

# Running the agent on desktop for 10 hours: first we fixed 6 bugs in our own monitor, then 3 P0s

> From July 31 to August 1 we ran the agent on desktop for nearly 10 hours straight, then spent a night using it "like a real developer". The biggest lesson was unexpected: during the soak we first found and fixed 6 flaws in our own monitoring code — one check returned "pass" whether or not the system had paused. Only then came three real P0s: shell calls hanging forever when the working directory vanished, a running task disappearing when you switched projects, and token counting starving the UI for 25 minutes.

Functional tests check "does it work"; soak tests check "does it still work after a long time". At the end of July we ran two rounds:

- **Round one (July 31)**: the desktop app ran for 9.66 hours straight, 1,134 samples, on DeepSeek V4 Flash, the session growing from 4 entries to 1,000, with timed probes sending commands, pausing and resuming, interrupting and interjecting;
- **Round two (the night of August 1)**: on throwaway target projects, doing what a normal developer does — take over a project, understand it, find bugs, fix them, review diffs, run tests, switch away to something else, come back and keep chatting, leave a long task running.

## Round one: the monitor was wrong in 6 places first

Partway through we realized some alerts were the monitor's own fault. One by one:

1. **"Stuck pending" judged by position** → judged by age. Parallel tool calls naturally finish out of order, so judging by position always misfires — 23 false alerts;
2. **"Stalled" judged by count** → by timestamp. Once entries hit the 1,000 memory cap the count stops changing — a permanent false alarm;
3. **Probe success judged by count** → same fix;
4. **Probes sent via simulated click** → direct injection. The input box was judged "outside the viewport", the click timed out, the command was never sent — and it was recorded as "no response". All three 90-second failures were this. After the fix the same probe passed in 1.5 seconds; a manual test took 2.9;
5. **Tool stats computed from in-memory entries** dropped when capped at 1,000 → take the run's peak and mark it as a lower bound;
6. **"Did it really pause?" judged by an unchanged count** → compare timestamps. With the count pinned at the cap, it was always equal — returning pass whether or not it paused, so the check had never actually worked.

The common lesson: **in a system with memory caps and virtualization, neither "counts" nor "visibility" can serve directly as criteria.** Alerts from before each fix were marked doubtful or void and excluded.

## Round one's real conclusion: one root cause behind every symptom

With the monitor noise gone, the picture was clear: after about 2 hours and 900-odd entries, the renderer's CPU saturated, which produced stuck pending entries (27 times), unresponsive commands and probe timeouts. When healthy, the command channel took 2.9 seconds — the channel itself was fine.

The saturation had two traceable sources:

- **An edge effect at the virtualization threshold.** At 926 entries the main timeline had 388 top-level rows — just under the "fully render below 400 rows" threshold — so it fully rendered and CPU saturated; past 400, virtualization kicked in and CPU recovered immediately. **The danger zone was exactly the band just below the threshold.**
- **Deep content hashing of the streaming row.** Profiling showed timeline row computation at 16% self time. The growing row changed on every event and its full text was hashed again and again — quadratic in text length per turn.

After a restart, probe latency fell from 76–458 seconds to 9–28 milliseconds and CPU from over 110% to under 4% — no irreversible damage.

## Round two: three P0s

Every criterion in round two rested on authoritative sources: the execution-state API, timeline pagination, React Profiler hooks, files on disk, running the tests ourselves. No guessing from page text, no guessing from counts.

### Shell calls hung forever when the working directory vanished

During the overnight run the app disappeared on its own; the logs showed an unhandled rejection: `spawn ENOTDIR`.

The way our process library fails when the working directory is gone made all three defenses miss (each verified):

| Defense | Reality |
|---|---|
| The surrounding try/catch | The spawn error isn't thrown synchronously — not caught |
| The child's error event | Never fires |
| The "don't reject on failure" option | Only covers non-zero exit codes, not spawn failure |

Only the promise rejection remained, and these code paths got results from events and never awaited the promise. Consequences: an unhandled rejection every time, filed into the crash archive (so real crashes found the quota full); worse, **the function never returned** and callers hung forever. Three execution paths were unguarded. After the fix: zero unhandled rejections, failures return an explicit exit code.

(A correction: we first thought this took the whole app down. We later confirmed the main process's unhandled-rejection handler only logs and doesn't exit; that exit came from the dev server exiting on its own, for reasons we didn't find. What's written above is the verified part.)

### Switch projects for a second, the running task vanishes

From the user's side: hand off a task, glance at another project's session, switch back — the UI is stuck on "thinking" forever, the stop button stays, and the last timeline entry is your own message. No answer, no error, no hint.

The chain: switching workspaces disposed the old project scope and built a new one; disposal tore down the connection without checking for running tasks or telling anyone, cutting that turn off. Switching back gave a brand-new runtime with run state starting from zero; but the UI's state mirror was global and persistent, and reconciliation only walked sessions in the authoritative list — this one was no longer there, so the ghost state was never corrected.

The fix: if a session is still running when you switch away, **keep it alive instead of tearing it down**; reuse it in place when you return to the same workspace; reclaim it after it goes idle. State queries can now ask about specific sessions and distinguish "it's idle" from "I don't know it". The verification script fails before the fix and passes after.

### The UI froze for 25 minutes while a long task wrote files

From the user's side: partway through a long task, the UI won't respond to clicks at all, though the task is still running in the background.

The data: renderer CPU above 110% for 25 minutes; a single read-only probe took 20–305 seconds; entries kept growing (data was arriving) but DOM nodes, mounted rows and scroll height **didn't change at all**; profiling showed constant element creation; React Profiler's commit callback **never fired once**.

Always rendering, never committing. **That's render starvation, not slow code.**

Root cause: for performance we had earlier moved the streaming token count from global state into a small dedicated store, shrinking the update scope from the whole tree to one `<span>`. But **the update frequency didn't change** — every token notified synchronously, dozens of times per second at peak. A small scope isn't a small cost: each notification is a synchronous React schedule, while the timeline's row mapping ran in an interruptible deferred render taking 40–50ms. The deferred render kept getting interrupted by the next token just before finishing, restarting each time, never reaching commit. The moment streaming stopped, CPU dropped from 110% to 4%.

The fix: batch token-count notifications per animation frame. Under the same load:

| | Before | After |
|---|---|---|
| Slowest main-thread probe | 14,907ms | 123ms |
| Commits during streaming | 0 | 24–61 per 10 seconds |

## Conclusion

In these two soak rounds we first found and corrected 6 flaws in the monitor's own criteria, all from using counts and visibility in a system with memory caps and virtualization; we switched to timestamps and authoritative APIs. With the monitor noise gone, round one traced the renderer's CPU saturation to two sources, and round two fixed three P0s: shell calls hanging forever when the working directory vanished, running tasks vanishing on project switch, and token counting starving the UI for 25 minutes. None of the three raised an error; they were found only through criteria built on authoritative data.
