Skip to content

Exact leaf sums and groups of any width - #6

Open
ukaratay wants to merge 11 commits into
mainfrom
dkaratay/opt-experiments
Open

ukaratay wants to merge 11 commits into
mainfrom
dkaratay/opt-experiments

Conversation

@ukaratay

@ukaratay ukaratay commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

predict() now returns the correctly rounded sum of the leaf values and accepts groups of any number of rows. The first four commits add the research kernels; the rest measure them, fold what paid into predict(), and remove the kernels again.

The native XGBoost baseline never predicted. Its config lacked missing and the iteration fields, so XGBoosterPredictFromDense returned -1 on every call ("Argument missing is required", XGBoost 3.2.0, checked 2026-10-03), and xgboost_native timed the error path: a flat ~180 µs at any model size. The released CSVs are flat the same way (289–295 µs Intel, 208–240 µs Arm, across all grid cells), so the XGBoost-native comparisons in paper/experiments/scripts/paper_numbers.py (lines 103, 116, 168–170) need a rerun. The harness now asserts the return code.

Leaves are summed exactly. Every leaf value of a model is an integer multiple of 2^-e for one scale e; a leaf adds v·2^e as an i128 to a difference array once per run of rows in its mask, and each row's prefix sum is rounded to f64 once. A leaf costs O(runs) instead of O(rows), and a prediction no longer depends on tree order, node layout, ablation mode or group width. A model with a nonfinite leaf, or with leaf exponents too far apart for 126 bits, adds in f64 in tree order instead; Forest::exact_sums() reports which.

Masks are u16, u32, u64, or Bits<W> of 64-bit words up to 1,024 rows. max_group_width can be any positive number, and wider groups run in 1,024-row pieces; a row predicted in a group of any width equals the row predicted alone, bit for bit. predict() also compiles out the ablation checks, which only predict_with_stats() reads.

On an M3 Pro (2026-10-03, 500 trees of depth 8), the new predict() is 1.11–1.13× faster than the previous f64 accumulation at 16 rows, 1.6–1.7× at 104–128 rows, and unchanged on Expedia sessions. Groups of 256–1,008 rows run 17–31% faster in one pass than in 128-row chunks. The run-list kernel was faster on monotone panels (1.2× at 256 rows, 2× at 1,008) but needs shape checks and a fallback path, so it is not included. Child-kind hints, 0–7% on a pinned Xeon, wait for the full suite run, which these numbers do not replace.

Breaking, so this needs a major version before publishing:

  • predict::MAX_GROUP_WIDTH is removed.
  • RowMask has new required methods and no u128 impl.
  • LightGBM predictions move in the last bits: up to 2.2e-15 on the four survival cells measured.

Tests:

  • A fixture test checks groups of 129–2,500 rows against single rows bit for bit; another checks that parse-time layouts cannot change a prediction.
  • The artifact suite's LightGBM tolerance against GTIL goes from 1e-14 to 1e-13. On the 1,008-step FLCHAIN panel, where margins reach ±500, GTIL's own f64 rounding is up to 1.06e-14 from the correctly rounded sum (math.fsum over the per-tree outputs, all 1,314,432 rows).
  • The ablation test used to call predict(), which ignores ablation. It now runs the ablations, skipping the monotone-scan modes on the tiled chunked-G cells, which break the monotonicity contract by construction.
  • test_binary_matches_json fails on the Expedia LightGBM artifact: Treelite's JSON dump writes "threshold": , for nonfinite thresholds, which the loader rejects as docs/treelite-loading.md documents.

@ukaratay
ukaratay added this pull request to stack #7 October 4, 2026 15:13
Base automatically changed from dkaratay/cargo-workspace to main October 6, 2026 23:16
ukaratay and others added 11 commits October 6, 2026 16:16
XGBoosterPredictFromDense rejected the old config (missing `missing` and
iteration_begin/iteration_end) and returned -1 on every call; the return code was
discarded, so xgboost_native timed the error path. Assert rc == 0 for the XGBoost
and LightGBM predict calls, and expose the last output for correctness checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012dbKy7QLkCLxqcaPjBhJZm
…elite output metadata

Port of the opt-experiments research kernels (src/opt.rs) and the layout override
hooks (parser::layout) onto rnewman/treelite-loading. Everything is gated by a
non-default `experimental` feature, so the default library, its public API and the
published crate are unchanged.

Adaptations to the new loading contract:
- Every kernel finalizes through Forest::finalize (divisor, base score, identity or
  sigmoid with alpha) instead of a hard-coded sigmoid; bit-identical to predict().
- Inline categorical evaluation (hint/catfree kernels) follows eval_split: f32 models
  round the category, negative and out-of-range categories are non-members.
- Input guards on all entry points (structural config, range, row capacity, buffer
  lengths); the explicit-stack kernel falls back to recursion past 64 pending
  entries, since trees may now be 256 edges deep.

tests/experimental.rs runs every bit-exact variant against predict() on the
independent import fixtures (all output modes, categories in both directions, f32,
boundary inputs, widths 32/64/128), exact accumulation against the GTIL oracle, and
the run-list kernel against the exact bitmask kernel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012dbKy7QLkCLxqcaPjBhJZm
opt_bench (paired, blocked comparison of kernel variants with exactness checks and
native XGBoost/LightGBM baselines), run_bench (run-list kernel vs chunked bitmask
kernels, phase knock-outs) and probe (single-kernel driver for cachegrind), behind
the benchmarks crate's `experimental` feature. Cells wider than the library's
128-row contract load with max_group_width clamped to 128 (wide::load_forest_any_width);
predict() is then chunked and the run-list kernel takes the full panel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012dbKy7QLkCLxqcaPjBhJZm
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012dbKy7QLkCLxqcaPjBhJZm
Format the one-line functions patch 1 added to external.rs. Under Rust 1.97.1
clippy rejected the research binaries and, on non-x86 targets, the opt module:
move the x86_64 cfg in pe_opt from the prefetch call to the whole prefetch
block, so `tgt` is not left unused elsewhere; split wide.rs's first doc
paragraph; and fix the pedantic and nursery lints in opt_bench, run_bench and
probe (semicolons, digit separators, `&mut` loops, clone_from, a redundant
clone, const fn, mul_add), allowing too_many_arguments and
many_single_char_names on the functions that need them.

Clippy now passes with -D warnings for every feature combination on
aarch64-apple-darwin, x86_64-unknown-linux-gnu (emeraldrapids and x86-64) and
aarch64-unknown-linux-gnu (neoverse-v2).
Groups were fixed-width (max_group_width), so CTR cells with variable-width
sessions failed the rows-multiple-of-width assertion. Use group_offsets.bin when
present, as sweep_bench does, and keep fixed-width groups otherwise.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
predict() always runs with the default AblationMode, yet partial_eval read the
flags from its context at every node, and predict_core filled a G-entry array
of row slices that only the ablation path's per-row partition reads. Resolve
the flags to the default when STATS is false, so the predict() instantiation
folds them away, and pass the per-row partition the group's rows and feature
count instead of the array. predict_with_stats still reads the flags at
runtime. On the pinned Xeon the opt kernel without these checks was 3-6%
faster than predict().

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
Every leaf value of a model is an integer multiple of 2^-e for one scale e
(src/exact.rs). predict() and predict_with_stats() now add a leaf's value
times 2^e to a difference array at the start and one past the end of each run
of rows in its mask (RowMask::run_starts, run_ends, for_each_bit) and convert
each row's prefix sum to f64 once. A leaf costs O(runs) instead of O(rows), and
a prediction is the correctly rounded sum of its leaf values, independent of
tree order, node layout and ablation mode. A model with a non-finite leaf, or
with leaf exponents too far apart for 126-bit sums, keeps f64 addition in tree
order; Forest::exact_sums() reports which path a model takes. predict_full()
still adds in f64 and can now differ from predict() in the last bits.

Tests: the full-walk comparisons use the f64 accumulation tolerance, and
ablation modes and parse-time layouts are checked bit for bit against
predict(). The ablation test used to call predict(), which ignores
config.ablation, so it never ran an ablation; it now calls
predict_with_stats() and skips the monotone-scan modes on cells whose data
break the declared monotonicity (the tiled chunked-G cells). The experimental
tests compare the f64-accumulating kernels with predict_full() and the exact
ones and run lists with predict() bit for bit.

On the M3 the new predict() is 1.1-1.2x faster at G=16 and 1.6-1.8x at
G=104-128 on the survival cells, unchanged on Expedia sessions, and
bit-identical to the opt kernels' exact accumulation.
The mask was u32, u64 or u128, chosen from max_group_width, which was limited
to 128. Masks are now u16 (up to 16 rows), u32, u64 and Bits<W> of W 64-bit
words (2, 4, 8 or 16 words, up to 1,024 rows), and max_group_width can be any
positive number. Groups wider than MAX_PIECE_ROWS (1,024) run in pieces of
that many rows, which also bounds the stack that masks passed down the
recursion use in deep trees. The pieces share the group's constant features,
and a row's sum does not depend on how rows are grouped, so a row predicted in
a group of any width equals the row predicted alone, bit for bit.

The workspace no longer keeps a 64-column copy of the group or a const G: the
threshold sweep gathers one feature at a time into a column buffer and sorts
(value, row) pairs instead of (value, mask) pairs, so its buffers grow with the
group rather than with the mask. The ablation-only partition functions read
values through a closure. MAX_GROUP_WIDTH is removed; u128 keeps its RowMask
impl for the opt kernels.

Tests: a fixture test predicts groups of 129 to 2,500 rows, with and without
ablation, against single-row predictions bit for bit, and the width sweep
covers the u16/u32 boundary. The artifact suite adds synthetic groups of 200,
1,000 and 2,500 rows and now loads the 256- and 1,008-step FLCHAIN panels,
which WalkerConfig used to reject. On the 1,008-step panel, where margins
reach +-500, GTIL's own f64 rounding is up to 1.06e-14 from the correctly
rounded sum (math.fsum over the per-tree outputs gives the same 1.06e-14 as
predict() over all 1,314,432 rows), so the LightGBM tolerance is now 1e-13.
The exact leaf built two full masks (run starts and run ends) and scanned both,
which touches every word of a Bits<16> mask even when the leaf's rows cover one
time range. RowMask::for_each_run replaces run_starts and run_ends: the integer
masks keep the two-shift form, and Bits<W> walks the set words only, extending
a run through full words, so a leaf costs O(runs plus the words they span).
Unit tests check every mask type against a bit-by-bit scan on random masks.

On the M3, predict() on a 1,008-row group goes from 0.75-0.81 to 0.69-0.72 of
the time of 128-row chunks; at 128 rows or fewer it is unchanged.
predict() now carries what measured worth keeping: exact leaf sums,
compile-time ablation, masks from u16 to Bits<16>, and groups of any width.
Remove the rest of the experimental feature: src/opt.rs (lockstep, prefetch,
the f64 difference array, the split and AVX-512 leaf adds, schedule-mask
caching, child-kind hints, the fall-through transform, compact nodes, huge
pages, the profiling kernels and the run-list kernel), the profile-guided and
hot/cold layout hooks in the parser, the opt_bench, run_bench and probe
binaries, wide.rs, tests/experimental.rs, the u128 mask and
RowMask::to_u128. The parser, both Cargo manifests and the benchmarks crate's
lib.rs are back to their state before the experiments, and the predict
helpers the kernels used are private again.

On monotone survival panels on the M3 the run-list kernel was 1.2x faster
than predict() at 256 rows and 2x at 1,008 rows, but it needs shape checks and
a fallback path; groups of any shape now run in one pass of Bits<W> masks,
17-31% faster than 128-row chunks. Child-kind hints measured 0-7% on the Xeon
and need a node field that pooled categorical splits already use; they can
come back with a measurement.

Two fixture tests cover what only the experimental tests or the artifact
suite checked: parse-time layouts give bit-identical predictions, and a model
without a common leaf scale adds in tree order, bit-identical to
predict_full().

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
@rnewman
rnewman force-pushed the dkaratay/opt-experiments branch from 1e28a39 to 28f4fcd Compare October 6, 2026 23:16
@ukaratay ukaratay self-assigned this Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant