Skip to content

feat(browser): pressure-based admission — brake the render limit on container CPU pressure, size claims to free capacity; v1.39.0 - #228

Draft
harper-joseph wants to merge 1 commit into
mainfrom
feat/pressure-admission
Draft

harper-joseph wants to merge 1 commit into
mainfrom
feat/pressure-admission

Conversation

@harper-joseph

@harper-joseph harper-joseph commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Adds an opt-in admission: { mode: 'pressure' } to the render worker. It does two things:

  1. Brake on CPU pressure. Each worker steps its concurrent-render limit between min and max according to the container's CPU pressure (cgroup v2 PSI).
  2. Size claims to free capacity. Each queue claim asks only for what can start soon, instead of a fixed batch.

fixed, the default, behaves exactly as before.

Why. Measured on a 20-pod production render fleet (15-core CFS limit per pod, 3 workers × 5 slots):

  • Overload costs output. Past ~95% pod CPU, output per pod falls and CPU per render rises. With the job mix held constant, 95–100% CPU gave −3% renders/pod/h and +7% CPU-seconds per render compared with 90–95%. On the 15-core pods, 14–16 cores gave 5,251 renders/pod/h against 5,597 at 12–14 cores.
  • Render cost varies ~20x by page (0.4 vs 8.5 CPU-s), so no single slot count is right all day.
  • Load is uneven across pods during a wave. At any minute, 12–43% of pods sat at ≥95% CPU with claimed jobs queued behind full slots, while 14–45% sat below 60%. Claiming a fixed batch of concurrency jobs lets a busy worker hold jobs an idle one could start.

For the human reviewer

  • No outside model family reviewed this. Codex and Gemini credentials aren't available. The planning review never ran; the pre-push review was a self-review plus one fresh-context Claude subagent. Both are the same model family as the author, so treat this as single-family coverage.
  • What the subagent review changed. The first version used a fixed claim size of 1, PSI avg10 as the signal, and a default max of 2 × concurrency. All three changed:
    • Claims are sized to free capacity (claimSize): free slots, or free prefetch-pool room. A fixed size of 1 capped each worker's start rate at one job per claim round trip, and a 1-job claim scans only 4 ready-set entries on the queue side.
    • Pressure is measured over each step's own interval, from the PSI some total= stall counter (pressureBetween). avg10 is a 10 s moving average, so a controller stepping every 5 s kept reacting to stale readings.
    • max defaults to concurrency (resolveAdmission), so by default the limit only brakes. Idle CPU with work waiting is also what a slow origin looks like. A limit that climbs on low pressure therefore sends more concurrent origin requests, and opens more Chrome pages, exactly then. Going above concurrency is an explicit operator choice, and the README says so.
  • Demand means "the limit is holding a job back" (jobHeldBack). The prefetch loop waits for a slot before it takes a job, so full slots count as demand there only while the pool holds a job. Pooled jobs with a slot free are prefetching ahead, not waiting on the limit.
  • Still open:
    • Prefetch pool depth: the pool can still hold up to prefetch.maxDepth (8) claimed jobs ahead per worker. It deepens and never shrinks, and this change doesn't touch it.
    • Claim rate: at saturation, claims become about one per job start (from one per 5 jobs). Queue-side cost per claim is small, but request volume to the queue port rises.
    • Thresholds: 15 / 30 were fitted to the 10 s average at steady state. The per-interval figure measures the same thing without the smoothing, but the thresholds haven't been re-fitted to it.
  • Not wired anywhere yet. render-service must pass admission through before this can be enabled. That is a follow-up PR there.

Changes

  • admission.ts (new): the AdmissionSettings type and the pure nextAdmissionLimit. It returns +1 under low pressure while a job is held back, −1 under high pressure, and floor(×0.75) above 2× high, clamped to [min, max]. With no reading, the limit holds.
  • util/cpu.ts: parseCpuStallUs / readCpuStallUs read the PSI stall counter, returning null without cgroup v2 PSI. pressureBetween turns two samples into a percent.
  • Worker.ts:
    • Setup: the constructor captures the admission settings. If PSI is missing it warns once at startup and the limit holds.
    • Waiting for a slot: awaitSlot waits while in-flight renders are at the limit, and wakes on a render finishing or on the limit rising.
    • Stepping: startAdmission starts at a random offset so the workers sharing a container step at different moments, and re-baselines at the first tick. samplePressure and stepAdmission do the per-tick reading and step.
    • Claims: the queue consumer now gets its claim size from claimSize.
    • Stats: the stats line reports the current limit as saturation.concurrency, plus saturation.admission {min, max, pressure}.
    • Browser page cap: the browser's maxActivePages follows max. Nothing enforces it; it only feeds the freeSlots stat.
  • RenderQueueConsumer.ts: takes an optional claimLimit getter that is read before each claim. The default is settings.jobClaimLimit, as before.
  • settings.ts:
    • New option: the admission option.
    • Claim ceiling: under pressure, jobClaimLimit defaults to max and acts as the ceiling on each claim.
    • Validation in resolveAdmission: min/max must be integers, 0 ≤ low < high < 100, and intervalMs must be within [1000, MAX_TIMER_MS]. A max below the default min pulls that default down to it.
    • Doc fix: corrects the stale jobClaimLimit doc, which said the default was concurrency * 2.
  • config.ts: exports MAX_TIMER_MS for that bound.
  • external/http.ts: the undici pool gets one connection per render that may post at once (max under pressure).
  • README.md: adds an admission section with defaults, the origin and memory caveat for raising max, and the measured thresholds. It also fixes the jobClaimLimit default in the options table.
  • package.json: 1.38.0 → 1.39.0.

Verification

  • test/admission.test.ts, 15 tests:
    • the nextAdmissionLimit step table;
    • stall-counter parsing and pressureBetween;
    • settings defaults and validation;
    • on a real RenderWorker with a stub browser, a gated renderer and a local queue endpoint:
      • a raised limit starts a waiting render immediately;
      • a lowered limit holds new starts;
      • fixed mode runs exactly concurrency renders;
      • demand with and without a prefetch pool, including a job still waiting across ticks;
      • claim sizing in both modes.
  • Full browser suite: 368/368 on the final head (596d0e6), plus the admission file 3× under 8 CPU-burning processes. Lint and prettier are clean.
  • Thresholds against production: sampled on the fleet at steady load, with a 15-core limit and 20 s windows. Per-pod pressure read 4 at 7 cores, 14 at 12, 19 at 13, 29–37 at 14.3 and 44–55 at 15. So 15 and 30 sit at about 80% and 95% of the pod's cores, and throughput peaked at 90–95%.
  • End to end: not observable until render-service exposes the option. Planned: enable it fleet-wide behind an env var (a per-pod A/B is invalid, because pods share one queue) and judge it over wave nights on pod CPU spread, renders/pod/h, wave duration and degraded-render rate.

🤖 Generated with Claude Code

…ontainer CPU pressure, size claims to free capacity; v1.39.0

Opt-in `admission: { mode: 'pressure' }`. Each worker steps its concurrent-render limit between
`min` and `max` from cgroup v2 PSI CPU pressure, measured over each step's own interval from the
`some total=` stall counter: up one while pressure is low and the limit is holding a job back, down
one while it is high, down a quarter when badly overloaded, on a staggered interval. `max` defaults
to `concurrency`, so by default the limit only brakes. A raised limit starts a waiting render at
once. Each queue claim is sized to what can start soon (free slots, or free prefetch-pool room) up to
`jobClaimLimit`. The default `fixed` mode is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a pressure-based admission control system that dynamically scales the concurrent render limit based on cgroup v2 CPU pressure (PSI) instead of relying on a fixed slot count. The changes span settings resolution, worker scheduling, connection pooling, and CPU stall monitoring utilities, backed by comprehensive unit tests. The feedback suggests a performance optimization to conditionally start the admission timer only when CPU pressure monitoring is actually available, preventing unnecessary timer overhead on unsupported environments like macOS or cgroup v1 hosts.

Comment on lines +259 to +265
if (this.lastStall === null) {
logger.warn(
{ limit: this.admissionLimit },
'pressure admission: no CPU pressure reading (needs cgroup v2 PSI) — the render limit holds'
);
}
this.startAdmission(this.admission);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

When CPU pressure monitoring (PSI) is unavailable (e.g., on macOS or cgroup v1 hosts), this.lastStall remains null. In this case, the admission limit will never change. To avoid the unnecessary overhead of starting the admission timer and repeatedly attempting to read the non-existent pressure file every few seconds, we can conditionally start the admission timer only when this.lastStall is not null.

Suggested change
if (this.lastStall === null) {
logger.warn(
{ limit: this.admissionLimit },
'pressure admission: no CPU pressure reading (needs cgroup v2 PSI) — the render limit holds'
);
}
this.startAdmission(this.admission);
if (this.lastStall === null) {
logger.warn(
{ limit: this.admissionLimit },
'pressure admission: no CPU pressure reading (needs cgroup v2 PSI) — the render limit holds'
);
} else {
this.startAdmission(this.admission);
}

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a dynamic, pressure-based admission control system for the prerender browser worker, allowing it to scale concurrent renders based on cgroup v2 CPU pressure rather than relying solely on a fixed concurrency limit. It includes the core admission logic, updates to the worker and queue consumer to support dynamic claim sizes, and extensive unit tests. The review feedback highlights a valid concern regarding a potential resource leak in startAdmission if the worker is destroyed while the initial startup timer is pending, suggesting a defensive check to prevent scheduling a new interval after shutdown.

Comment on lines +441 to +450
this.admissionTimer = globalThis.setTimeout(
() => {
// Re-baseline rather than step: the window since construction is an arbitrary slice of an
// interval, and each step should read one whole interval.
this.samplePressure();
this.admissionTimer = setInterval(step, admission.intervalMs);
this.admissionTimer.unref();
},
Math.floor(Math.random() * admission.intervalMs)
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

When starting the admission timer, if destroy() or shutdown() is called while the initial setTimeout is pending, there is a risk of starting an un-cleared setInterval if the callback executes. Although destroy() attempts to clear the timer, adding a defensive check for this.shuttingDown (or this.admissionTimer === null) inside the setTimeout callback ensures that we do not schedule a new setInterval during or after shutdown, preventing potential resource leaks.

Suggested change
this.admissionTimer = globalThis.setTimeout(
() => {
// Re-baseline rather than step: the window since construction is an arbitrary slice of an
// interval, and each step should read one whole interval.
this.samplePressure();
this.admissionTimer = setInterval(step, admission.intervalMs);
this.admissionTimer.unref();
},
Math.floor(Math.random() * admission.intervalMs)
);
this.admissionTimer = globalThis.setTimeout(
() => {
if (this.shuttingDown) return;
// Re-baseline rather than step: the window since construction is an arbitrary slice of an
// interval, and each step should read one whole interval.
this.samplePressure();
this.admissionTimer = setInterval(step, admission.intervalMs);
this.admissionTimer.unref();
},
Math.floor(Math.random() * admission.intervalMs)
);

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant