Repository navigation
Conversation
XGBoosterPredictFromDense rejected the old config (missing `missing` and iteration_begin/iteration_end) and returned -1 on every call; the return code was discarded, so xgboost_native timed the error path. Assert rc == 0 for the XGBoost and LightGBM predict calls, and expose the last output for correctness checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012dbKy7QLkCLxqcaPjBhJZm
…elite output metadata Port of the opt-experiments research kernels (src/opt.rs) and the layout override hooks (parser::layout) onto rnewman/treelite-loading. Everything is gated by a non-default `experimental` feature, so the default library, its public API and the published crate are unchanged. Adaptations to the new loading contract: - Every kernel finalizes through Forest::finalize (divisor, base score, identity or sigmoid with alpha) instead of a hard-coded sigmoid; bit-identical to predict(). - Inline categorical evaluation (hint/catfree kernels) follows eval_split: f32 models round the category, negative and out-of-range categories are non-members. - Input guards on all entry points (structural config, range, row capacity, buffer lengths); the explicit-stack kernel falls back to recursion past 64 pending entries, since trees may now be 256 edges deep. tests/experimental.rs runs every bit-exact variant against predict() on the independent import fixtures (all output modes, categories in both directions, f32, boundary inputs, widths 32/64/128), exact accumulation against the GTIL oracle, and the run-list kernel against the exact bitmask kernel. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012dbKy7QLkCLxqcaPjBhJZm
opt_bench (paired, blocked comparison of kernel variants with exactness checks and native XGBoost/LightGBM baselines), run_bench (run-list kernel vs chunked bitmask kernels, phase knock-outs) and probe (single-kernel driver for cachegrind), behind the benchmarks crate's `experimental` feature. Cells wider than the library's 128-row contract load with max_group_width clamped to 128 (wide::load_forest_any_width); predict() is then chunked and the run-list kernel takes the full panel. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012dbKy7QLkCLxqcaPjBhJZm
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012dbKy7QLkCLxqcaPjBhJZm
Format the one-line functions patch 1 added to external.rs. Under Rust 1.97.1 clippy rejected the research binaries and, on non-x86 targets, the opt module: move the x86_64 cfg in pe_opt from the prefetch call to the whole prefetch block, so `tgt` is not left unused elsewhere; split wide.rs's first doc paragraph; and fix the pedantic and nursery lints in opt_bench, run_bench and probe (semicolons, digit separators, `&mut` loops, clone_from, a redundant clone, const fn, mul_add), allowing too_many_arguments and many_single_char_names on the functions that need them. Clippy now passes with -D warnings for every feature combination on aarch64-apple-darwin, x86_64-unknown-linux-gnu (emeraldrapids and x86-64) and aarch64-unknown-linux-gnu (neoverse-v2).
Groups were fixed-width (max_group_width), so CTR cells with variable-width sessions failed the rows-multiple-of-width assertion. Use group_offsets.bin when present, as sweep_bench does, and keep fixed-width groups otherwise. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
predict() always runs with the default AblationMode, yet partial_eval read the flags from its context at every node, and predict_core filled a G-entry array of row slices that only the ablation path's per-row partition reads. Resolve the flags to the default when STATS is false, so the predict() instantiation folds them away, and pass the per-row partition the group's rows and feature count instead of the array. predict_with_stats still reads the flags at runtime. On the pinned Xeon the opt kernel without these checks was 3-6% faster than predict(). Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
Every leaf value of a model is an integer multiple of 2^-e for one scale e (src/exact.rs). predict() and predict_with_stats() now add a leaf's value times 2^e to a difference array at the start and one past the end of each run of rows in its mask (RowMask::run_starts, run_ends, for_each_bit) and convert each row's prefix sum to f64 once. A leaf costs O(runs) instead of O(rows), and a prediction is the correctly rounded sum of its leaf values, independent of tree order, node layout and ablation mode. A model with a non-finite leaf, or with leaf exponents too far apart for 126-bit sums, keeps f64 addition in tree order; Forest::exact_sums() reports which path a model takes. predict_full() still adds in f64 and can now differ from predict() in the last bits. Tests: the full-walk comparisons use the f64 accumulation tolerance, and ablation modes and parse-time layouts are checked bit for bit against predict(). The ablation test used to call predict(), which ignores config.ablation, so it never ran an ablation; it now calls predict_with_stats() and skips the monotone-scan modes on cells whose data break the declared monotonicity (the tiled chunked-G cells). The experimental tests compare the f64-accumulating kernels with predict_full() and the exact ones and run lists with predict() bit for bit. On the M3 the new predict() is 1.1-1.2x faster at G=16 and 1.6-1.8x at G=104-128 on the survival cells, unchanged on Expedia sessions, and bit-identical to the opt kernels' exact accumulation.
The mask was u32, u64 or u128, chosen from max_group_width, which was limited to 128. Masks are now u16 (up to 16 rows), u32, u64 and Bits<W> of W 64-bit words (2, 4, 8 or 16 words, up to 1,024 rows), and max_group_width can be any positive number. Groups wider than MAX_PIECE_ROWS (1,024) run in pieces of that many rows, which also bounds the stack that masks passed down the recursion use in deep trees. The pieces share the group's constant features, and a row's sum does not depend on how rows are grouped, so a row predicted in a group of any width equals the row predicted alone, bit for bit. The workspace no longer keeps a 64-column copy of the group or a const G: the threshold sweep gathers one feature at a time into a column buffer and sorts (value, row) pairs instead of (value, mask) pairs, so its buffers grow with the group rather than with the mask. The ablation-only partition functions read values through a closure. MAX_GROUP_WIDTH is removed; u128 keeps its RowMask impl for the opt kernels. Tests: a fixture test predicts groups of 129 to 2,500 rows, with and without ablation, against single-row predictions bit for bit, and the width sweep covers the u16/u32 boundary. The artifact suite adds synthetic groups of 200, 1,000 and 2,500 rows and now loads the 256- and 1,008-step FLCHAIN panels, which WalkerConfig used to reject. On the 1,008-step panel, where margins reach +-500, GTIL's own f64 rounding is up to 1.06e-14 from the correctly rounded sum (math.fsum over the per-tree outputs gives the same 1.06e-14 as predict() over all 1,314,432 rows), so the LightGBM tolerance is now 1e-13.
The exact leaf built two full masks (run starts and run ends) and scanned both, which touches every word of a Bits<16> mask even when the leaf's rows cover one time range. RowMask::for_each_run replaces run_starts and run_ends: the integer masks keep the two-shift form, and Bits<W> walks the set words only, extending a run through full words, so a leaf costs O(runs plus the words they span). Unit tests check every mask type against a bit-by-bit scan on random masks. On the M3, predict() on a 1,008-row group goes from 0.75-0.81 to 0.69-0.72 of the time of 128-row chunks; at 128 rows or fewer it is unchanged.
predict() now carries what measured worth keeping: exact leaf sums, compile-time ablation, masks from u16 to Bits<16>, and groups of any width. Remove the rest of the experimental feature: src/opt.rs (lockstep, prefetch, the f64 difference array, the split and AVX-512 leaf adds, schedule-mask caching, child-kind hints, the fall-through transform, compact nodes, huge pages, the profiling kernels and the run-list kernel), the profile-guided and hot/cold layout hooks in the parser, the opt_bench, run_bench and probe binaries, wide.rs, tests/experimental.rs, the u128 mask and RowMask::to_u128. The parser, both Cargo manifests and the benchmarks crate's lib.rs are back to their state before the experiments, and the predict helpers the kernels used are private again. On monotone survival panels on the M3 the run-list kernel was 1.2x faster than predict() at 256 rows and 2x at 1,008 rows, but it needs shape checks and a fallback path; groups of any shape now run in one pass of Bits<W> masks, 17-31% faster than 128-row chunks. Child-kind hints measured 0-7% on the Xeon and need a node field that pooled categorical splits already use; they can come back with a measurement. Two fixture tests cover what only the experimental tests or the artifact suite checked: parse-time layouts give bit-identical predictions, and a model without a common leaf scale adds in tree order, bit-identical to predict_full(). Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
rnewman
force-pushed
the
dkaratay/opt-experiments
branch
from
October 6, 2026 23:16
1e28a39 to
28f4fcd
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
predict()now returns the correctly rounded sum of the leaf values and accepts groups of any number of rows. The first four commits add the research kernels; the rest measure them, fold what paid intopredict(), and remove the kernels again.The native XGBoost baseline never predicted. Its config lacked
missingand the iteration fields, soXGBoosterPredictFromDensereturned -1 on every call ("Argumentmissingis required", XGBoost 3.2.0, checked 2026-10-03), andxgboost_nativetimed the error path: a flat ~180 µs at any model size. The released CSVs are flat the same way (289–295 µs Intel, 208–240 µs Arm, across all grid cells), so the XGBoost-native comparisons inpaper/experiments/scripts/paper_numbers.py(lines 103, 116, 168–170) need a rerun. The harness now asserts the return code.Leaves are summed exactly. Every leaf value of a model is an integer multiple of 2^-e for one scale e; a leaf adds v·2^e as an
i128to a difference array once per run of rows in its mask, and each row's prefix sum is rounded to f64 once. A leaf costs O(runs) instead of O(rows), and a prediction no longer depends on tree order, node layout, ablation mode or group width. A model with a nonfinite leaf, or with leaf exponents too far apart for 126 bits, adds in f64 in tree order instead;Forest::exact_sums()reports which.Masks are
u16,u32,u64, orBits<W>of 64-bit words up to 1,024 rows.max_group_widthcan be any positive number, and wider groups run in 1,024-row pieces; a row predicted in a group of any width equals the row predicted alone, bit for bit.predict()also compiles out the ablation checks, which onlypredict_with_stats()reads.On an M3 Pro (2026-10-03, 500 trees of depth 8), the new
predict()is 1.11–1.13× faster than the previous f64 accumulation at 16 rows, 1.6–1.7× at 104–128 rows, and unchanged on Expedia sessions. Groups of 256–1,008 rows run 17–31% faster in one pass than in 128-row chunks. The run-list kernel was faster on monotone panels (1.2× at 256 rows, 2× at 1,008) but needs shape checks and a fallback path, so it is not included. Child-kind hints, 0–7% on a pinned Xeon, wait for the full suite run, which these numbers do not replace.Breaking, so this needs a major version before publishing:
predict::MAX_GROUP_WIDTHis removed.RowMaskhas new required methods and nou128impl.Tests:
math.fsumover the per-tree outputs, all 1,314,432 rows).predict(), which ignores ablation. It now runs the ablations, skipping the monotone-scan modes on the tiled chunked-G cells, which break the monotonicity contract by construction.test_binary_matches_jsonfails on the Expedia LightGBM artifact: Treelite's JSON dump writes"threshold": ,for nonfinite thresholds, which the loader rejects asdocs/treelite-loading.mddocuments.