Skip to content

predict: three kernel changes from the kernel study, gated on both VMs - #11

Open
ukaratay wants to merge 3 commits into
dkaratay/v2-expfrom
dkaratay/pr-kernel
Open

ukaratay wants to merge 3 commits into
dkaratay/v2-expfrom
dkaratay/pr-kernel

Conversation

@ukaratay

@ukaratay ukaratay commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Three local changes to production predict, from the profile and assembly study of the hot kernel paths. Outputs are bit-identical, and the stack passed the VM kernel gate on both machines (2026-10-05): faster on every gated cell under fixed function alignment, G=1 by 10-13% on Intel and 24-27% on Arm. The library only: src/predict/ and one new test.

  • The first walk in the tree loop. Every tree began with a call to partial_eval, even when its walk over constant splits reached a leaf before any varying split: every tree of a G=1 group, and most trees of small models. eval_tree, inlined into the tree loop, walks the constant prefix itself and adds a leaf it reaches over all rows as one run. Anything else continues in partial_eval, unchanged and in the same tree order, so every row still adds one leaf per tree in tree order. On the M4 Pro (ABBA, 3 rounds, default | 64-byte function alignment): G=1 -17.0 | -17.0%, single-row serving -17.3 | -16.4%, T=500 L=4 -3.1 | -5.0%; Expedia T=500 L=8 +1.1 | +0.6% over 5 quiet rounds. The work counters are identical for every runtime and load variant on 12 cells.
  • One clear of the difference array. Each piece zeroed prefix_starts and, after finalization, the whole difference array. The first fill did nothing, and finalization now zeroes each entry as it reads it. A new test reuses one predictor across rejected calls and groups of alternating widths and constants, on a forest whose prefix group bails at levels 0, 1 and 2, and checks every result bit for bit against a fresh predictor; removing either clear fails it. On the M4 Pro: G=1 -1.2 | -1.1%, single-row serving -1.6 | -1.7%, against an A/A band of ±1%.
  • A prefetch of each next tree's start node, in models of 4,096 nodes or more, where the node pool outgrows the L1 data cache (64 KiB). An ungated prefetch cost G=1 2-3%, so smaller models run the loop without it. On the M4 Pro it ranges from -5.4 | -1.9% (Expedia T=500 L=8) to +0.5 | -1.4% (T=1000 L=16 LightGBM); it is a memory-side change, so the VM gate decides.

The gate is the one in infra/ in the last PR of this stack: TreeWalker production built from the base (A) and from the candidate (B), timed alone in alternating processes, A B B A, three times per alignment, with default function alignment and with every function aligned to 64 bytes, on the kernel study's 10 cells. A slowdown counts only if it persists under fixed alignment. TreeWalker's between-process spread there is under 0.6% on Intel and 1.2% on Arm. The gate ran on these same patches (2026-10-05), cherry-picked here onto #10.

Checked at every commit: cargo fmt --check, clippy with and without external-bench, cargo test, the external-bench library tests, ruff, ruff format, mypy and pytest. At the head, the README's cross-target clippy matrix, 8 combinations (aarch64-apple-darwin, x86_64 at emeraldrapids and x86-64, aarch64 Linux at neoverse-v2; with and without external-bench, since pmu arrives with the runner).

ukaratay and others added 3 commits October 7, 2026 08:09
…ee loop.

Every tree began with a call to partial_eval, even when its walk over
constant splits reached a leaf before any varying split: every tree of a
G=1 group, and most trees of small models. The kernel study found
single-row serving at G=1 about 20% slower than the full walk, mostly that
per-call overhead.

eval_tree, inlined into the tree loop, walks the constant prefix itself
and adds a leaf reached that way over all rows as one run (diff[0] and
diff[n]). Anything else continues in partial_eval, unchanged and in the
same tree order, so every row still adds one leaf per tree in tree order.
The leaf addition moves into add_leaf, shared by both. Outputs are
bit-identical, and the work counters are identical for every runtime and
load variant, on 12 cells (8 real, 4 fixtures).

On the M4 Pro (ABBA, 3 rounds, default | 64-byte function alignment):
G=1 -17.0 | -17.0%, single-row serving -17.3 | -16.4%, T=500 L=4
-3.1 | -5.0%, T=1000 L=16 XGBoost -1.7 | -1.3%; Expedia T=500 L=8
+1.1 | +0.6% over 5 quiet rounds.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
…dant fill.

Each piece zeroed prefix_starts and, after finalization, the whole
difference array. Trees outside every prefix group keep the zero start
the workspace was created with, and grouped trees are rewritten for every
piece, so the first fill did nothing. Finalization now zeroes each entry
as it reads it, and the last one. Nothing between the first leaf addition
and finalization returns early, so no call leaves the array dirty.

A new test reuses one predictor across rejected calls and groups of
alternating widths and constants, on a forest whose prefix group bails at
levels 0, 1 and 2, and checks every result bit for bit against a fresh
predictor; removing either clear fails it. On the M4 Pro: G=1
-1.2 | -1.1%, single-row serving -1.6 | -1.7% (default | fixed
alignment), against an A/A band of ±1%; neutral elsewhere.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
…s or more.

A tree's first node is a dependent load at the start of its walk. In
models whose node pool outgrows the L1 data cache (4,096 nodes are
64 KiB), the tree loop now prefetches the next tree's start node, at its
prefix-group start. Smaller models run the loop without it, since an
ungated prefetch cost G=1 2-3%. A prefetch changes no result.

On the M4 Pro, on its own against the base (default | fixed alignment):
T=500 L=8 -2.6 | -3.4%, Expedia T=500 L=8 -5.4 | -1.9%, T=1000 L=16
XGBoost -1.2 | -1.4%, credit what-if -1.2 | -0.9%, T=1000 L=16
LightGBM +0.5 | -1.4%. It is a memory-side change, so the VMs decide.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
@ukaratay
ukaratay added this pull request to stack #7 October 7, 2026 16:38
@ukaratay ukaratay self-assigned this Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant