Skip to content

[staging] ci/warmer-nl-toolchain - #30

Closed
b7r6 wants to merge 41 commits into
mainfrom
ci/warmer-nl-toolchain
Closed

b7r6 wants to merge 41 commits into
mainfrom
ci/warmer-nl-toolchain

Conversation

@b7r6

@b7r6 b7r6 commented Oct 3, 2026

Copy link
Copy Markdown
Owner

Fork-internal staging PR to exercise CI before the upstream PR opens. Not for merge.

amankrx and others added 30 commits September 29, 2026 03:59
* Admin API lists connected workers and queued demand

* Docs: the admin listings

* Admin demand listing reads the queue the way the matcher does

* Admin demand listing carries the client's last keepalive

* Admin demand listing leaves out a record that moved on or went away between reads
…perties cached per pass (TraceMachina#2805)

* Add a least_loaded allocation strategy so a newly joined worker takes the next action

* best_fit placement with a priority preference, typed properties cached per pass
…unreadable row skipped, a wait for the new index (TraceMachina#2803)

* Redis: read a cursor page in the RESP3 shape the first page already parses

* Scheduler: the Redis backend gets the retention default, so completed records expire

* Scheduler: the listing skips a record it cannot read instead of ending

* Redis: wait for a new search index to finish its scan before reading it
…tion, one store read per queued action (TraceMachina#2806)

* Give room that opens during a matching pass to the oldest waiting action

* Serve a listed action from the record the listing already read

* Matching pass: parked actions keep the listing order, a higher memory report is room opening, the parked metric counts confirmed sends
…tdown guard that waits for the drain (TraceMachina#2807)

* GoingAway with a drain flag so a shutting-down worker stops receiving work, and reconnect backoff with jitter

* Count the original shutdown guard so a SIGTERM waits for the worker to drain
* Dispatch is acknowledged or declined by the worker

* Dispatch ack: the answer is recorded before the liveness refresh, a late decline does not pause, counters count requeues, a dead channel is a decline

* Dispatch ack: the timeout counts from the scheduler clock at the send, and a late acknowledgement is answered with a kill

* Keep the pause through a decline, name the memory property, defer kills on a full channel
…ancelled (TraceMachina#2809)

* Kill the whole process group when an action times out or is cancelled

* Run the process-group kill test without namespaces

* Write the kill test's pid file under the action root

* Kill the process group from the child wrapper so every kill path covers it

* Use the imported error and log paths in the group kill helper
…d orphaned action directories (TraceMachina#2810)

* Bound worker output capture, uploads, precondition scripts and orphan directories

* Detach the orphan sweep task so it outlives its spawn handle

* Orphan sweep: carry on past a directory it cannot remove, skip symlinks, and log a quiet sweep at debug

* Cut log excerpts by bytes and take the cleanup mark in the orphan sweep

* Bound open output files with one semaphore and spill outside the file budget

* Gate the gated store to unix with the test that uses it
…uard before inputs are fetched (TraceMachina#2812)

* Worker: kill an action that exceeds its memory reservation before the pod's cgroup limit does

* Kill on a single sample at twice the memory ceiling

* Worker honours the platform properties the scheduler reserved for the action

* Worker refuses an action whose disk reservation exceeds the free space before fetching its inputs

* Use the child wrapper's group kill for the memory ceiling

* Log a ceiling breach whose action already ended instead of dropping the send result

* Shorten the ceiling-breach log line so both formatters agree

* Enforce each axis alone, sample Pss, count admitted disk reservations, overlay properties only with enforcement on

* Gate the unix-only test helpers with the tests that use them
…hina#2820)

* Worker advertises CPU and memory from its own cgroup limits

* Advertise CPU in whole cores by default, millicores only when told

* Headroom follows resource_enforcement unless set; refuse to start without a cgroup and without the properties
* Worker readiness follows scheduler registration

* Regenerate main's config reference and have the snippet lint read it too

* A lost registration reads Initializing, and a readiness path equal to the status path is refused
…on, and resource usage with an outcome (TraceMachina#2822)

* Worker: SIGTERM grace, output kept on kill, per-action TMPDIR, timeout clamp, usage v2 with outcome

* Read the timed-out result's error from the action result in the grace test

* Borrow the finished result's error in the grace test

* Drop the wall time assertion from the memory kill test

* Namespaces: pass SIGTERM through the stub so a namespaced action gets its kill grace

* Follow the split enforcement axes and the regenerated bindings

* Use the namespace helpers as main has them and a store key for the kept-output check

* Gate the exit-signal read to unix and drop the marker removal result
…lers no longer hold its output open (TraceMachina#2823)

* Worker reaps the zombies actions leave behind, and an action's stragglers no longer hold its output open

* Take the action's group id before the guard owns the child, and match the zombie sets by lookup

* Reaper: own the children we spawn, scan stat off the runtime, forget reaped pids; drain bounded only for killed or straggling actions; biased read

* Take the reaper's last sweep by reference
…und over /tmp (TraceMachina#2824)

* Give each action a private /tmp: its own tmp directory bound over /tmp inside the mount namespace

* Say temporary directory in prose where vale reads tmp as a misspelling

* Private /tmp: no mask under a bound /tmp, the size cap only with isolate_tmp, a probe that runs the real sequence, isolate_tmp judged against the resolved mode

* Map the gid in the namespace probe as the action path does
…writes (TraceMachina#2825)

* Worker: a disk soft limit and peak_disk_kb from the files an action writes

* Walk the action directory through the crate's blocking spawn, and expect the sample to catch the write in progress

* Give the integration test suite the tempfile dependency the disk walk test uses

* Disk limit: the action's own files by modification time, the walk as a polled task, periodic walks only for the soft limit, the disk figure counts as sampled

* Pass the walk start to the Linux sampler and give the disk guard test slack

* Time the disk walk by the filesystem's clock
…orker as the last step (TraceMachina#2829)

* Scheduler: memory escalation with retry budgets by cause, the whole worker as the last step

* Escalation: no veto on a whole-worker ask, max_kb caps the ceiling, draining workers left out, percent validated, a zero base refused, a max_steps test, the shared requeue metric
…ion environment, namespaces, measured usage, and a wait at the cap (TraceMachina#2830)

* Persistent workers: a configurable pool with an idle sweeper, the action environment, namespaces, measured usage, and a wait at the cap

* Start the persistent worker sampler with no ceilings

* Persistent worker tests pass PATH to the worker process

* Persistent worker processes get the worker's additional environment too

* Keep the unix cfg on the shell echo script

* Persistent workers: key prints no environment values, the wait at the cap is raced against the kill and bounded by the action's time, CPU as a delta, one environment helper, the pool on by default

* Windows persistent worker tests pass the whole environment
…h a cold-start policy, and where each reservation came from (TraceMachina#2832)

* Historical resource scheduler: hints by digest keys, size classes with a cold-start policy, and where each reservation came from

* Write the eternal test timeout in hours

* Hints: every key family a record names, the ladder guarded by what the action carries, zero hints fall to cold start, rules and property types checked at load, disk on the cold start
…aceMachina#2843)

* Worker: build the cleanup mark's guard after the lock is released, so a second mark cannot deadlock the runtime on its own mutex

* Worker: an action's cleanup waits for the mark instead of stepping aside, and the sweep and a retry's stale removal hold it
…EOF (TraceMachina#2844)

* Store: a read that stops short of the digest's size is an error, not EOF, in the Redis store and the verify store

* Store: DataLoss outranks the inner send error on an over-long read, and a copy that fails size verification is dropped from the caches below
…rectory (TraceMachina#2845)

* Worker: read capacity and free memory from the worker's own cgroup directory, not the mount root

* Worker: the cgroup limit comes from the nearest limited ancestor, no limit anywhere keeps the typed properties, and the failure names the cgroup
…l reach it (TraceMachina#2846)

* Worker: an input fetch that never completes fails the action at max_download_timeout, and a kill reaches it, instead of holding its slot

* Worker: a kill token every phase watches, so a kill ends a fetch or an upload; cleanup retries the removal; the tests clean up
…ina#2811)

* Add opt-in generic gRPC event-sink transport

Support string-keyed write-only event stores without changing CAS or Action Cache semantics. Carry the actual payload size in ByteStream resources, validate configuration, and regenerate the configuration reference.

Refresh the change onto main and retain the declaration-order lint fix. Replaces the same-author PR commits f775a29 and d5b1e4c.

AI-assisted implementation.

* Validate event-sink acknowledgements and exercise lost-ack retries

Reject event-sink responses whose committed size differs from the full payload. Cover incomplete acknowledgements and retries after the receiver consumes an event but loses its response. Document durable receiver acknowledgements, stable event-key deduplication, and the limits of publisher recovery.

Validated the 13 gRPC-store and 11 configuration tests with Bazel, including Clippy and rustfmt, and ran the documentation and repository hooks.

AI-assisted implementation at Marcus Eagan's direction.
…hina#2826)

should_timeout_operation ended an Executing action once Action.timeout
had passed since it was assigned. The worker enforces the same timeout,
but its clock starts when the command starts, after input fetch, so the
scheduler always fired first. The action was requeued onto the worker
still running it, which refused with AlreadyExists; the retries ran out,
and when the worker's own DEADLINE_EXCEEDED result arrived it was
treated as a stray and the worker was disconnected, taking every other
action on it down too.

The scheduler now enforces Action.timeout in every liveness case, at
Action.timeout plus no_event_action_timeout, measured from the later of
the assignment and the worker's last update on the action. A live
worker reports first, so no liveness special-casing is needed, and the
scheduler still ends an action whose worker never reports: a hung input
fetch, an unkillable child, a broken timeout_handled_externally wrapper,
or an orphan no instance reaps.

The client's DEADLINE_EXCEEDED message now names the deadline that
fired, instead of always citing no_event_action_timeout.

This is the interim: the scheduler cannot observe when the command
starts. TraceMachina#2827 tracks the worker reporting command start as an operation
transition, so the deadline can be anchored to that event.

Co-authored-by: Johnny <johnny2@flightlines.ai>
…store, and a missing input names its digest so the client re-uploads it (TraceMachina#2855)
…the worker evicted and idle (TraceMachina#2856)

* Worker: no scheduler call awaited in the run loop, a deadline on keepalives and acknowledgements that reconnects, and HTTP/2 keepalive on the scheduler channel

* Worker: the GoingAway calls on shutdown carry the same deadline

* Worker: the layout CI's rustfmt asks for

* Keep the scheduler stream flowing: ordered processing off the reader, no deadlines on stream messages

* Qualify Code in the worker API server test

* Box the action future before the launch closure

* Box the action chain inside the action closure

* Raise the binary's recursion limit for the worker's action future

* Use try_update for the filesystem store generation counter
amankrx and others added 10 commits October 2, 2026 01:18
…a#2859)

Co-authored-by: config reference bot <bot@tracemachina.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
* chore(deps): update rust crate serial_test to v4
* Update MODULE.bazel.lock
…oolchain identity

The Nix / Bazel Dev lane gets 0/2723 remote cache hits. It compiles inside
`nix develop`, where the default C++/Rust toolchain resolves to nix-store
paths (/nix/store/<hash>-clang|gcc|rust) that are baked into every action
key, so its digests never match the cache the Native lanes warm off a
non-nix toolchain (the Bazel 9 Native lane gets 2273/3915 against the same
endpoint). If the remote cache warms one Bazel config it should warm both;
this closes that gap by keying the lane on a stable, relocatable toolchain
identity rather than host paths.

- Factor the hermetic LLVM platform + downloaded Rust toolchains that
  nl-rbe already uses into a new, endpoint-free `nl-toolchain` config, and
  have nl-rbe inherit it via `--config=nl-toolchain`. These toolchains key
  on downloaded CONTENT, not host paths, so they produce identical action
  digests on a bare runner, inside `nix develop`, or on a remote worker.
  No behavior change for nl-rbe (same flags, now shared).
- Apply `--config=nl-toolchain` to the Bazel Dev lane. This is cache-read
  only (no remote executor added): the lane keeps running locally, it just
  keys its compiles against the hermetic identity so they can hit a warm
  cache.

Verification:
- PROVEN (measured): action keys are byte-identical. A `bazel aquery` proof
  run compares three ways and shows the nix+nl-toolchain command lines match
  the bare-runner nl-toolchain command lines and no longer reference
  /nix/store:
  https://github.com/b7r6/nativelink/actions/runs/37059855343
- PROJECTED (not yet measured): with a warm cache the lane should drop from
  ~670s to ~250s. This is a projection -- proven here is key alignment, not
  a full warm read.
- REQUIRES ONE UPSTREAM CHANGE to realize the win: the trusted main-push
  cache warmer must also build with `--config=nl-toolchain`, so it warms the
  same key this lane now requests (today's Native warmer uses bare
  `--extra_toolchains` without the `@llvm` platform pin). One line; happy to
  do it in a follow-up once this lands.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…y (--config=nl-toolchain)

Realizes the warm-cache contract TraceMachina#2874 sets up. That change keys the
Bazel Dev (nix) lane's reads on the endpoint-free nl-toolchain config —
the hermetic LLVM platform + downloaded Rust toolchains that nl-rbe
already uses — because toolchains keyed on downloaded content produce
identical action digests on a bare runner, inside nix develop, or on a
remote worker. But nobody writes that digest space yet: the trusted
main-push Bazel 9 lane (the only holder of a write key) warms with bare
--extra_toolchains, host-toolchain keys the Dev lane can never hit.

Switch the Bazel 9 lane from --extra_toolchains=@rust_toolchains//:all
to --config=nl-toolchain (which contains that same line plus the @llvm
platform pin). On main pushes this lane becomes the warmer for the
hermetic digest space; on PRs it reads from it. The Bazel 8 lane is
untouched (separate, lockfile_mode=off digest space).

Transition cost, stated honestly: the first runs after this merges are
cold for everyone until the first main push warms the new key space —
one build. Measured so far is key ALIGNMENT, not the warm win: the
aquery proof run shows nix+nl-toolchain command lines byte-identical to
bare-runner nl-toolchain with no /nix/store references
(https://github.com/b7r6/nativelink/actions/runs/37059855343). The warm
hit-rate itself cannot be measured from a fork (no write credentials);
projected from the Bazel 9 lane's current warm behavior, the Dev lane
drops from ~11 min to roughly the Native lane's ~5 min.

Stacked on TraceMachina#2874 (requires its nl-toolchain config). Builds on the
remote-cache lanes and @palfrey's TraceMachina#2864 disk-cache handling.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Some of the pull request description still needs filling in:

  • What and why is missing. Please keep the template's headings.
  • How was this verified? is missing. Please keep the template's headings.
  • Risk is missing. Please keep the template's headings.
  • AI assistance is missing. Please keep the template's headings.

Edit the description and this check re-runs on its own. The sections exist because they are the parts a reviewer cannot get from the diff: why the change is needed, how you know it works, what breaks if it is wrong, and which AI tools helped ("None" is a complete answer).

…y (--config=nl-toolchain)

Realizes the warm-cache contract TraceMachina#2874 sets up. That change keys the
Bazel Dev (nix) lane's reads on the endpoint-free nl-toolchain config —
the hermetic LLVM platform + downloaded Rust toolchains that nl-rbe
already uses — because toolchains keyed on downloaded content produce
identical action digests on a bare runner, inside nix develop, or on a
remote worker. But nobody writes that digest space yet: the trusted
main-push Bazel 9 lane (the only holder of a write key) warms with bare
--extra_toolchains, host-toolchain keys the Dev lane can never hit.

Switch the Bazel 9 lane from --extra_toolchains=@rust_toolchains//:all
to --config=nl-toolchain (which contains that same line plus the @llvm
platform pin). On main pushes this lane becomes the warmer for the
hermetic digest space; on PRs it reads from it. The Bazel 8 lane is
untouched (separate, lockfile_mode=off digest space).

Transition cost, stated honestly: the first runs after this merges are
cold for everyone until the first main push warms the new key space —
one build. Measured so far is key ALIGNMENT, not the warm win: the
aquery proof run shows nix+nl-toolchain command lines byte-identical to
bare-runner nl-toolchain with no /nix/store references
(https://github.com/b7r6/nativelink/actions/runs/37059855343). The warm
hit-rate itself cannot be measured from a fork (no write credentials);
projected from the Bazel 9 lane's current warm behavior, the Dev lane
drops from ~11 min to roughly the Native lane's ~5 min.

Stacked on TraceMachina#2874 (requires its nl-toolchain config). Builds on the
remote-cache lanes and @palfrey's TraceMachina#2864 disk-cache handling.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@b7r6

b7r6 commented Oct 7, 2026

Copy link
Copy Markdown
Owner Author

Closing: this CI experiment concluded — its findings shipped upstream via the TraceMachina#2874/TraceMachina#2877/TraceMachina#2880/TraceMachina#2886/TraceMachina#2897 series. Branch preserved.

@b7r6 b7r6 closed this Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants