Skip to content

treewalker-exp: one package for data preparation, with manifests and new workloads - #10

Open
ukaratay wants to merge 10 commits into
dkaratay/api-v2from
dkaratay/v2-exp
Open

ukaratay wants to merge 10 commits into
dkaratay/api-v2from
dkaratay/v2-exp

Conversation

@ukaratay

@ukaratay ukaratay commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

The experiment scripts become one package, treewalker-exp, and data preparation is rebuilt around manifests, so a rerun can no longer mix old models with new data. The first commit moves paper/experiments/ to experiments/ and benchmarks/ to experiments/benchmarks/, renames only.

The package reproduces the released preparation byte for byte. Today's scripts, run twice from a clean checkout of #9, produced the same 30 files both times; the package's output matches them for a SUPPORT panel, an Expedia session cell and a credit what-if cell, on the old library versions. The check is pinned at the refactor commit, because the workloads commit after it changes some of those files on purpose.

  • Manifests. Every model writes model.json and every cell cell.json. A model is keyed by dataset, split, training parameters, library versions and the training data's hash; a cell by its model, generator, parameters and data. Prep reuses one only when its key matches and every file has its recorded hash. PREP_POLICY versions the status rules, so changing them re-evaluates old cells on resume without retraining. A compiled baseline is reused only for the same model, settings, compiler path and version, target, host CPU, and tl2cgen or llvmlite version.
  • Workloads, 600 models and 1,094 factorial cells:
    • FLCHAIN horizons 256, 512 and 1,024;
    • what-if on all four datasets: raw covariates drawn jointly from one training entity, derived features recomputed, draws coupled across G and k;
    • a ranking size curve over the 3,439 test sessions with at least 32 candidates, where size n is the first n candidates;
    • panels declare interaction_x0_t increasing, since x0 is age.
  • Expedia changes. prop_country_id is declared varying: it is constant in 99.44% of searches, and the old prep dropped the 2,236 cross-border ones so that a constant declaration held. Candidates are now in a seeded random order per session. The file was sorted by prop_id, which the T=500, L=8 LightGBM model splits on in 10.0% of its 86,753 splits, so TreeWalker's exact sums saw fewer, longer runs: on the M4 Pro (2026-10-04) the sorted order was 1.3% faster with LightGBM in each of 8 alternating rounds, and a median 1.4% with XGBoost.
  • Dependencies: lightgbm 4.7.0, xgboost 3.4.1 (xgboost-cpu on Linux, which drops CUDA and NCCL), treelite 4.7.2 and the rest at their current releases. tl2cgen moves into uv.lock, and compile.py runs from its own committed script lock with today's lleaves settings. PyPI is the project's first index, so a Shopify user config with the package proxy no longer rewrites the locks. Treelite 4.7.2 fixed its macOS libomp rpath, so the DYLD_LIBRARY_PATH workaround is gone.
  • Fixtures are regenerated with Treelite 4.7.2, numpy 2.5.3, scikit-learn 1.9.1 and scipy 1.18.1. Only the 17 binaries changed, each in one header byte, and tests/import.rs passes unchanged.

The most expensive cell, FLCHAIN horizon 1,024 at T=2000, L=16, trained in 948 s with 14.3 GB peak memory on the M4 Pro (2026-10-04), for 8.2 million nodes, within every parser limit. Its Treelite JSON would have been 2.53 GB, which no loader accepts; JSON is now written only for the models the tests read. On that model correctness.rs misses its 1e-13 bound against GTIL at 1.21e-13. TreeWalker's own f64 full walk misses by more (2.17e-13), so this is GTIL's sequential rounding over 2,000 deep trees. The next PR's exact fsum oracle replaces that comparison.

Infra: git_ref no longer defaults to neurips2026, which lacks this package, and startup uses the new commands. prepare_chunked.py stays until the runner rewrite removes its callers. The root crate's include patterns are anchored, so cargo package no longer picks up 26 license and README files from a root .venv.

@ukaratay
ukaratay added this pull request to stack #7 October 5, 2026 01:35
ukaratay and others added 7 commits October 6, 2026 16:16
…hem.

paper/ held only paper/experiments, so the experiment tooling now sits at
experiments/, and the benchmark crate, which exists only to run those
experiments, moves to experiments/benchmarks. Python packaging in the next
commits builds on this layout.

The commit only renames and updates paths: the workspace member and the
crate's path dependency, the artifact directories the benchmark tests build
from CARGO_MANIFEST_DIR, the scripts that counted parent directories to
find the repository, .gitignore, the README, the docs and the infra
scripts. rustfmt joins three Rust lines whose strings got shorter. The
crate keeps its name and binaries, and cargo package still lists only the
library.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
A pattern without a slash matches at any depth, so "LICENSE",
"README.md", "Cargo.toml" and "CHANGELOG.md" also packaged the matching
files of every wheel in a root .venv, which uv sync creates. The patterns
with a slash were already relative to the root; all of them now start with
one, so the list reads the same way throughout. cargo package lists the
same 79 files with or without a .venv: the library's sources and tests, the
two loading docs, the manifests, README, CHANGELOG and LICENSE.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
…age.

The data preparation lived in four scripts around a 1,558-line utils.py,
671 lines of which no script reached. It is now one package,
experiments/treewalker_exp, built with uv_build and run as
`uv run treewalker-exp`. It is a repository tool rather than a wheel: it
finds the checkout by walking up to experiments/grids.toml, or takes
--repo, so it also runs from outside the repository root.

- grids.toml holds the grids; the package resolves them into models and
  cells. The factorial suite covers what `prepare.py --grid all` built,
  the ablation suite its anchors, and scenario-v1 the released credit
  cells (whatif-v1, in their released layout).
- formats has the one reader and writer of each artifact file, so
  prepare_chunked.py and audit_f32.py now read test_data.bin through it.
- paths keeps every path; pooch fetches SUPPORT, FLCHAIN and credit into
  experiments/data/raw/ against pinned SHA-256s, instead of a temporary
  directory with no checks.
- compile.py runs lleaves as a subprocess under its own script lock,
  compile.py.lock, with a request that carries the released recipe:
  fblocksize 34, llc -O3, fp-contract on, the output path and the worker
  count. Its own defaults (O2, fp-contract fast) used to differ from the
  in-process call. LLVM tools come from the request or PATH.
- Each compiled baseline records a compile identity next to the library:
  the model's hash, the settings, the compiler's resolved path, version
  and target triple, the host CPU (and for lleaves the CPU that llc's
  -mcpu=native resolves to), and tl2cgen's version or compile.py and its
  lock. A library is reused only when that identity and its hash match,
  so one built by another compiler or on another machine is rebuilt. llvmlite stays below 0.48, the last LLVM 20 release,
  because the VMs' llc is LLVM 20. The unused PGO, BOLT and opt paths,
  with their Homebrew and /usr/local paths, are gone.
- tl2cgen moves into a locked baselines dependency group, and lleaves
  out of the project lock. The aarch64 <cstdint> compiler wrapper is a
  tracked script, infra/scripts/cxx-cstdint, instead of a file written to
  /usr/local/bin at boot.
- ofat and the fixed and geometric Expedia groupings are not carried
  over. prepare_chunked.py stays until the runner redesign replaces it.

The VM startup script and the README call the new commands. Every uv run
there passes --group baselines, because uv run removes packages outside
the groups it syncs. Terraform's git_ref loses its neurips2026 default,
since that tag predates everything the startup script now runs.

The project and compile.py declare PyPI as a uv index. A project index
takes priority over indexes in a user's uv configuration, so uv sync
--locked, uv lock and uv lock --check leave the locks unchanged on a
machine whose configuration adds another index.

prepare produces byte-identical artifacts to the base commit's scripts
(on the same library versions) for SUPPORT nt500_md8_h16, Expedia
nt500_md8 and the whatif-v1 credit cell k4_G16: all 30 files match a
reference that the old scripts rebuilt in a clean directory, twice, with
identical results.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
…ouches for.

Prep used to reuse any file that existed, so a rerun could mix old models
with new data or library versions, and nothing recorded what a cell was.
Each model now writes model.json and each cell cell.json. A model is keyed
by dataset, source hash, split, training parameters, library versions and
the training data's hash; a cell by its model, generator, parameters and
data hashes. Prep reuses one only when its key matches and every file has
its recorded hash, and rebuilds anything else. A retrained model drops its
compiled baselines, which the released runner would otherwise load.
compile-baselines records each library's request, effective settings and
hash in the model's cells.

`treewalker-exp manifest`, also run by prepare, resolves a suite into
artifacts/manifests/<suite>.json: every model, workload and cell with its
key and status, so the runner needs no grid logic. Released files keep
their places, so today's sweep_bench still reads the panel and session
cells.

New workloads, about 1,100 factorial cells in all:
- FLCHAIN horizons 256, 512 and 1,024 at T in {500, 2000} and L in {8, 16}.
- ranking-sessions-v2: each test session's candidates in one seeded
  random order. The file is sorted by prop_id within each search, and the
  Expedia LightGBM model at T=500, L=8 splits on prop_id in 10.0% of its
  86,753 splits, so rows that a prop_id split sends the same way were
  contiguous, so the rows reaching a leaf formed fewer, longer runs, and
  exact sums update once per run. Requests do not arrive sorted by property. Measured on the M4 Pro (2026-10-04), with
  production predict_groups over the same 10,000 test sessions: the seeded
  order takes 1.3% longer than the sorted one with LightGBM, in each of 8
  alternating rounds, and a median 1.4% with XGBoost, where 3 of 8 rounds
  went the other way. Outputs are bit-identical once rows are mapped back.
- ranking-cohort-v2: the test sessions with at least 32 candidates. The
  size-n subset (4, 8, 16, 32) is each session's first n candidates in that
  seeded order, so the subsets nest and keep one order across sizes, because
  row order decides how exact sums group rows into runs.
- panel-v2 declares interaction_x0_t increasing when x0 >= 0: x0 is age in
  SUPPORT and FLCHAIN, fixed within a patient, while the time fraction
  increases.
- whatif-v2 on all four datasets: the model's most-split eligible raw
  covariates, values drawn jointly from one training entity per scenario,
  and draws coupled across G and k. Survival donors are patients, not
  expanded rows, which favor patients who live longer; the time step is
  one seeded draw per entity, and interaction_x0_t is recomputed per row.
  Each cell records the requested, selected and realized k. It is labelled
  a targeted stress test; credit keeps its frozen monetary pool.

Expedia's prop_country_id, the property's country, is now declared
varying. It is the same for every candidate of 99.44% of searches, but
2,236 of the 399,344 searches with 2-128 candidates span countries, and the
session filter used to drop them so that a constant declaration held. Only
the ten search-level features are filtered on now, so the sampled sessions
change, and with them the Expedia models.

Prep checks each cell's contracts before writing its references: constant
features equal within every group, declared monotonic features monotonic
over their non-missing values. A violation fails the cell. Each model
records TreeWalker's exact import limits, 32,767 nodes per tree and a
32-million-node pool, and a model over them is marked unsupported. Each
cell records an upper bound on its varying predicates (categorical splits
are not deduplicated) and only warns above the parser's 65,535; the
runner's preflight, which loads the model, decides.

The Treelite JSON dump is written only for the models that the Rust tests
and the byte-identical check read, listed in grids.toml, and never above
the JSON loader's 64 MiB: nothing else reads it, since the runner and
tl2cgen load the binary and QuickScorer the native text. The pilot's
LightGBM model (FLCHAIN horizon 1,024, T=2000, L=16) would have written
2.53 GB. The dump is derived from the binary and tracked apart from the
model's files, so a change to the list exports or deletes it on the next
run and never retrains.

PREP_POLICY versions the rules behind every derived record: a model's
limits and JSON export, a cell's status, contracts and warnings. Cell keys
include it and model.json records it, so changing the rules re-evaluates
every model and cell on resume. A model is not retrained, and a cell whose
data and model are unchanged keeps its references; only its status and
records are recomputed.

Tests cover the generators' invariants (nesting, the shared cohort,
the session order, coupled draws, the recomputed interaction and its
monotonic declaration, realized k, the feature roles, the contracts), the resume rules (a matching manifest is reused; changed data,
library versions, models or file hashes rebuild), the limit boundaries,
predicate deduplication, the JSON allowlist, and resuming from manifests
in the two earlier formats, stored as fixtures: a cell that the first
version marked unsupported by an off-by-one predicate rule becomes ready
without --force, retraining or new references, and a stale JSON dump over
the limit is deleted. The byte-identical check passes at the previous
commit, where the 30 released files match; this one changes the panel
configurations and the Expedia sessions and models on purpose.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
Python 3.14 evaluates annotations lazily (PEP 649), so the import does
nothing in the seven scripts that kept it. Each script still runs:
paper_numbers.py reports 91 checks with no mismatches, and plot.py,
gen_heatmap_tex.py, decomposition_validation.py, summarize_scenario.py,
audit_f32.py and prepare_chunked.py run as before.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
The lower bounds move to lightgbm 4.7.0, xgboost 3.4.1, treelite 4.7.2,
polars 1.44.2, numpy 2.5.3, pyarrow 25.0.1, plotnine 0.15.8, requests
2.34.2 and h5py 3.16.0, and uv.lock is relocked against PyPI, which also
takes scipy to 1.18.1. This comes after the byte-identical check because
newer libraries can train different models.

- XGBoost and Treelite move together: Treelite 4.7.0 cannot import DART
  models from XGBoost 3.3 or later (fixed in 4.7.1). XGBoost stays at
  3.4.1, the latest on PyPI.
- On Linux the dependency is xgboost-cpu: the xgboost wheel there is a
  57.6 MB CUDA build that pulls in NCCL, while xgboost-cpu is 5.8 MB and
  has no CUDA dependency. It has no macOS wheels, so macOS keeps xgboost.
- Treelite 4.7.1 fixed libtreelite's rpath, so Treelite now loads libomp
  from Homebrew without DYLD_LIBRARY_PATH; the last instruction to set it
  goes.

On SUPPORT nt500_md8_h16, Expedia nt500_md8 and the credit cells, data,
LightGBM trees, Treelite JSON and references are unchanged. The XGBoost
models change (XGBoost 3.3 rewrote the hist quantile sketch). The LightGBM
text gains one parameter line, [gpu_device_id_list: ], and the Treelite
binary's header records 4.7.2. The
artifact correctness suite passes on these models: within 1.3e-15 for
LightGBM and 6.4e-7 for XGBoost.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
The fixture generator pinned only its direct dependencies, on the command
line. It now declares them as inline script metadata (Treelite 4.7.2,
NumPy 2.5.3, scikit-learn 1.9.1, SciPy 1.18.1) with a committed script
lock, generate.py.lock, and runs with uv run --locked --script. Like the
project, the script declares PyPI as its uv index. Treelite
4.7.2 loads libomp without DYLD_LIBRARY_PATH, so that instruction goes,
and the manifest's oracle comes from the installed Treelite.

Every fixture and reference is regenerated. Only the binary exports
change, each in one byte: the Treelite patch version in the header, 0 to
2. The JSON dumps, the GTIL and sklearn references and the inputs are
byte-identical, and manifest.json records the new versions with Treelite
GTIL 4.7.2 as the oracle. tests/import.rs passes unchanged.
docs/treelite-loading.md moves its four 4.7.0 references to 4.7.2.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
@rnewman
rnewman self-requested a review October 7, 2026 00:08

@rnewman rnewman left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  1. [P1] The dependency upgrade breaks VM startup. pyproject.toml:18 installs XGBoost 3.4.1, but startup.sh still builds 3.2.0 and copies its library over the installed package. XGBoost rejects this version mismatch during import, so both VMs stop before preparation. Align the native build versions with the lockfile; LightGBM has the same discrepancy.

  2. [P2] Execution manifests mark stale cells ready. manifest.py:47 trusts existing cell documents without validating their files. I prepared both frameworks, changed the split seed, and rebuilt only LightGBM. The shared test data changed, invalidating XGBoost’s recorded hash, but the execution manifest still reported both cells ready. Validate artifacts before carrying forward ready.

  3. [P2] Workload filtering silently drops overlapping cells. grids.py:133 retains only the first workload name when deduplicating cells. Consequently, prepare --suite factorial --workload whatif-credit-full --dry-run selects 14 cells instead of 32: the other 18 belong to whatif-credit-core after deduplication. Preserve workload memberships or filter before deduplicating.

ukaratay and others added 3 commits October 7, 2026 08:02
The dependency upgrade moved uv.lock to LightGBM 4.7.0 and XGBoost 3.4.1,
but the startup script still built 4.6.0 and 3.2.0 from source and copied
them over the installed packages' libraries. XGBoost refuses a library of
another version at import, so both VMs stopped before preparation. The
native builds now use the locked versions.

Found by Richard Newman in review.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
The execution manifest carried every cell's recorded status forward
without checking its files. Cells share files, such as a model's test
data: with both frameworks prepared, changing the split seed and
rebuilding only LightGBM rewrote the shared test data, so XGBoost's
recorded hash no longer matched, yet the manifest still reported both
cells ready. A ready cell whose files no longer have their recorded
hashes is now `stale`; each shared file is hashed once per manifest.

Found by Richard Newman in review.
…loads share.

A suite keeps one cell per ID, named for the first workload that lists
it, so `--workload` missed shared cells: `prepare --suite factorial
--workload whatif-credit-full --dry-run` selected 14 cells instead of 32,
because the other 18 also belong to whatif-credit-core. A suite now
records every workload naming each cell, and the filter matches any of
them. Each cell keeps the first workload's name, so no cell's identity
changes: the factorial still resolves to the same 600 models and 1,094
cells.

Found by Richard Newman in review.

Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants