Skip to content

Remote compute-node workflow - #241

Draft
BenWibking wants to merge 2 commits into
AMReX-Codes:mainfrom
BenWibking:feature/server-memory-cache-budget
Draft

BenWibking wants to merge 2 commits into
AMReX-Codes:mainfrom
BenWibking:feature/server-memory-cache-budget

Conversation

@BenWibking

Copy link
Copy Markdown
Contributor

A remote server currently starts each dataset's block cache with the client's budget (normally 1 GiB), regardless of the memory allocated to its Slurm job. This adds opt-in --cache-memory-fraction 0.5: the server detects its memory allowance after srun launches it and assigns half to cached payloads shared across all datasets and connections.

Linux detection resolves the process's cgroup v1/v2 mount and checks visible ancestor limits, combines those with available Slurm per-node/per-CPU memory information, and bounds the result by physical RAM. A Slurm job with no discoverable limit fails startup instead of assuming it owns the whole node. The detected source and resulting allowance are logged to stderr without disturbing the stdio protocol.

Block and volume-grid caches share admission accounting and reclaim unpinned entries under pressure. Payloads remain charged while requests or escaped shared references retain them, including after dataset closure. Initial block budgets use the detected allowance; subsequent client requests are capped by it. The existing per-dataset volume-grid cap still applies. Omitting the option preserves existing behavior.

Includes a new Slurm workflow guide with a named-allocation launcher, a Riker example, automatic-sizing instructions, cleanup, and troubleshooting, linked from the README and bundled help.

Validation:

  • macOS Release builds of the server and Qt client, with compiler warnings treated as errors.
  • 17 focused CTest cases passed, including fixtures, memory detection, shared-cache concurrency and allocation-failure rollback, cross-connection reclamation, combined block/grid accounting, the real server over stdio, and existing local query/pipeline regressions.
  • Standalone AddressSanitizer/UndefinedBehaviorSanitizer runs passed for the cache and memory-detection tests.
  • Bash example syntax, local documentation/resource paths, and whitespace checks passed.

The allowance bounds cached payloads, not process RSS: temporary reads, rendering buffers, particles, and other job processes still require headroom. Detection runs once at startup. Linux hierarchy cases are tested through injected OS inputs; a live Slurm/Riker allocation has not been tested locally.

@BenWibking BenWibking changed the title Auto-size a shared server cache from job memory limits Remote compute-node workflow Sep 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant