Repository navigation
Remote compute-node workflow - #241
Draft
BenWibking wants to merge 2 commits into
Draft
BenWibking wants to merge 2 commits into
BenWibking wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A remote server currently starts each dataset's block cache with the client's budget (normally 1 GiB), regardless of the memory allocated to its Slurm job. This adds opt-in
--cache-memory-fraction 0.5: the server detects its memory allowance aftersrunlaunches it and assigns half to cached payloads shared across all datasets and connections.Linux detection resolves the process's cgroup v1/v2 mount and checks visible ancestor limits, combines those with available Slurm per-node/per-CPU memory information, and bounds the result by physical RAM. A Slurm job with no discoverable limit fails startup instead of assuming it owns the whole node. The detected source and resulting allowance are logged to stderr without disturbing the stdio protocol.
Block and volume-grid caches share admission accounting and reclaim unpinned entries under pressure. Payloads remain charged while requests or escaped shared references retain them, including after dataset closure. Initial block budgets use the detected allowance; subsequent client requests are capped by it. The existing per-dataset volume-grid cap still applies. Omitting the option preserves existing behavior.
Includes a new Slurm workflow guide with a named-allocation launcher, a Riker example, automatic-sizing instructions, cleanup, and troubleshooting, linked from the README and bundled help.
Validation:
The allowance bounds cached payloads, not process RSS: temporary reads, rendering buffers, particles, and other job processes still require headroom. Detection runs once at startup. Linux hierarchy cases are tested through injected OS inputs; a live Slurm/Riker allocation has not been tested locally.