Repository navigation
Conversation
…ee loop. Every tree began with a call to partial_eval, even when its walk over constant splits reached a leaf before any varying split: every tree of a G=1 group, and most trees of small models. The kernel study found single-row serving at G=1 about 20% slower than the full walk, mostly that per-call overhead. eval_tree, inlined into the tree loop, walks the constant prefix itself and adds a leaf reached that way over all rows as one run (diff[0] and diff[n]). Anything else continues in partial_eval, unchanged and in the same tree order, so every row still adds one leaf per tree in tree order. The leaf addition moves into add_leaf, shared by both. Outputs are bit-identical, and the work counters are identical for every runtime and load variant, on 12 cells (8 real, 4 fixtures). On the M4 Pro (ABBA, 3 rounds, default | 64-byte function alignment): G=1 -17.0 | -17.0%, single-row serving -17.3 | -16.4%, T=500 L=4 -3.1 | -5.0%, T=1000 L=16 XGBoost -1.7 | -1.3%; Expedia T=500 L=8 +1.1 | +0.6% over 5 quiet rounds. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
…dant fill. Each piece zeroed prefix_starts and, after finalization, the whole difference array. Trees outside every prefix group keep the zero start the workspace was created with, and grouped trees are rewritten for every piece, so the first fill did nothing. Finalization now zeroes each entry as it reads it, and the last one. Nothing between the first leaf addition and finalization returns early, so no call leaves the array dirty. A new test reuses one predictor across rejected calls and groups of alternating widths and constants, on a forest whose prefix group bails at levels 0, 1 and 2, and checks every result bit for bit against a fresh predictor; removing either clear fails it. On the M4 Pro: G=1 -1.2 | -1.1%, single-row serving -1.6 | -1.7% (default | fixed alignment), against an A/A band of ±1%; neutral elsewhere. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
…s or more. A tree's first node is a dependent load at the start of its walk. In models whose node pool outgrows the L1 data cache (4,096 nodes are 64 KiB), the tree loop now prefetches the next tree's start node, at its prefix-group start. Smaller models run the loop without it, since an ungated prefetch cost G=1 2-3%. A prefetch changes no result. On the M4 Pro, on its own against the base (default | fixed alignment): T=500 L=8 -2.6 | -3.4%, Expedia T=500 L=8 -5.4 | -1.9%, T=1000 L=16 XGBoost -1.2 | -1.4%, credit what-if -1.2 | -0.9%, T=1000 L=16 LightGBM +0.5 | -1.4%. It is a memory-side change, so the VMs decide. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
ukaratay
added this pull request to stack #7
October 7, 2026 16:38
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three local changes to production predict, from the profile and assembly study of the hot kernel paths. Outputs are bit-identical, and the stack passed the VM kernel gate on both machines (2026-10-05): faster on every gated cell under fixed function alignment, G=1 by 10-13% on Intel and 24-27% on Arm. The library only:
src/predict/and one new test.partial_eval, even when its walk over constant splits reached a leaf before any varying split: every tree of a G=1 group, and most trees of small models.eval_tree, inlined into the tree loop, walks the constant prefix itself and adds a leaf it reaches over all rows as one run. Anything else continues inpartial_eval, unchanged and in the same tree order, so every row still adds one leaf per tree in tree order. On the M4 Pro (ABBA, 3 rounds, default | 64-byte function alignment): G=1 -17.0 | -17.0%, single-row serving -17.3 | -16.4%, T=500 L=4 -3.1 | -5.0%; Expedia T=500 L=8 +1.1 | +0.6% over 5 quiet rounds. The work counters are identical for every runtime and load variant on 12 cells.prefix_startsand, after finalization, the whole difference array. The first fill did nothing, and finalization now zeroes each entry as it reads it. A new test reuses one predictor across rejected calls and groups of alternating widths and constants, on a forest whose prefix group bails at levels 0, 1 and 2, and checks every result bit for bit against a fresh predictor; removing either clear fails it. On the M4 Pro: G=1 -1.2 | -1.1%, single-row serving -1.6 | -1.7%, against an A/A band of ±1%.The gate is the one in
infra/in the last PR of this stack: TreeWalker production built from the base (A) and from the candidate (B), timed alone in alternating processes, A B B A, three times per alignment, with default function alignment and with every function aligned to 64 bytes, on the kernel study's 10 cells. A slowdown counts only if it persists under fixed alignment. TreeWalker's between-process spread there is under 0.6% on Intel and 1.2% on Arm. The gate ran on these same patches (2026-10-05), cherry-picked here onto #10.Checked at every commit:
cargo fmt --check, clippy with and withoutexternal-bench,cargo test, the external-bench library tests, ruff, ruff format, mypy and pytest. At the head, the README's cross-target clippy matrix, 8 combinations (aarch64-apple-darwin, x86_64 at emeraldrapids and x86-64, aarch64 Linux at neoverse-v2; with and withoutexternal-bench, sincepmuarrives with the runner).