Repository navigation
fix(cloud): a runaway job on a cloud machine is stopped before the machine stops answering - #197
Merged
Merged
Conversation
…chine stops answering Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
andrewcai8
enabled auto-merge (squash)
October 8, 2026 08:02
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A cloud chat's machine locked up twice on 2026-10-07: the agent's own test runs filled its 8 GB, the machine thrashed instead of the kernel killing anything, and envd stopped answering, so the chat could not be reached until it was rebooted. Telling agents to use worker copies does not prevent it: a census of 11 live chats found 9 had never used them.
The machine now runs earlyoom, the standard userspace OOM killer: when free memory falls to 8% it stops the biggest process, preferring test runners and builds and avoiding envd, systemd, sshd, the agent CLIs and the T3 server. It is set up next to the envd memory drop-in on every box setup and wake. When earlyoom already runs with these settings the step does nothing (0.17 s live); otherwise install and setup start as a transient systemd unit, so a slow apt never delays a wake (0.8 s live, earlyoom active about 7 s later), and the next wake retries if it failed.
Measured on this template (32 bun workers holding 350 MB each over a 1.5 GB file corpus):
Upstream's open pingdotgg#14338 (per-workload systemd scopes for systemd-oomd on desktops) was considered: these machines have no PSI, which systemd-oomd needs, and earlyoom needs neither PSI nor a user systemd manager.
Claude Opus 5.5, Claude Code
🤖 Generated with Claude Code