Skip to content

fix(cloud): a runaway job on a cloud machine is stopped before the machine stops answering - #197

Merged
andrewcai8 merged 1 commit into
mainfrom
fix/cloud-earlyoom
Oct 8, 2026
Merged

andrewcai8 merged 1 commit into
mainfrom
fix/cloud-earlyoom

Conversation

@andrewcai8

Copy link
Copy Markdown
Owner

A cloud chat's machine locked up twice on 2026-10-07: the agent's own test runs filled its 8 GB, the machine thrashed instead of the kernel killing anything, and envd stopped answering, so the chat could not be reached until it was rebooted. Telling agents to use worker copies does not prevent it: a census of 11 live chats found 9 had never used them.

The machine now runs earlyoom, the standard userspace OOM killer: when free memory falls to 8% it stops the biggest process, preferring test runners and builds and avoiding envd, systemd, sshd, the agent CLIs and the T3 server. It is set up next to the envd memory drop-in on every box setup and wake. When earlyoom already runs with these settings the step does nothing (0.17 s live); otherwise install and setup start as a transient systemd unit, so a slow apt never delays a wake (0.8 s live, earlyoom active about 7 s later), and the next wake retries if it failed.

Measured on this template (32 bun workers holding 350 MB each over a 1.5 GB file corpus):

  • Without it: envd missed checks for about 30 s and answered in up to 9 s.
  • With a lower memory cap on envd's user cgroup instead: no misses, answers up to 2.9 s.
  • With earlyoom: no misses, answers within 1 s; it stopped 14 bun workers, and stand-ins for the T3 server (node) and claude kept running; the kernel OOM killer never fired.

Upstream's open pingdotgg#14338 (per-workload systemd scopes for systemd-oomd on desktops) was considered: these machines have no PSI, which systemd-oomd needs, and earlyoom needs neither PSI nor a user systemd manager.

Claude Opus 5.5, Claude Code

🤖 Generated with Claude Code

…chine stops answering

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@andrewcai8
andrewcai8 enabled auto-merge (squash) October 8, 2026 08:02
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Oct 8, 2026
@andrewcai8
andrewcai8 merged commit 9e2a007 into main Oct 8, 2026
25 checks passed
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 5.0 KiB 5.0 KiB 0 B (0.0%) 6.8 KiB ✅
Codex Thread snapshot wire 3.8 KiB 3.8 KiB 0 B (0.0%) 4.9 KiB ✅
Codex Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Codex Live turn WebSocket decoded 20.9 KiB 20.9 KiB 0 B (0.0%) 29.3 KiB ✅
Codex Live turn messages 2 2 0 (0.0%) 8 ✅
Claude Total thread wire 5.0 KiB 5.0 KiB 0 B (0.0%) 6.8 KiB ✅
Claude Thread snapshot wire 3.8 KiB 3.8 KiB 0 B (0.0%) 4.9 KiB ✅
Claude Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Claude Live turn WebSocket decoded 21.2 KiB 21.2 KiB 0 B (0.0%) 29.3 KiB ✅
Claude Live turn messages 1 1 0 (0.0%) 8 ✅

Baseline: d677a5a · PR result: e903649 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 108.5 KiB
  • Claude decoded thread snapshot: 108.8 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant