Repository navigation
Conversation
…hem. paper/ held only paper/experiments, so the experiment tooling now sits at experiments/, and the benchmark crate, which exists only to run those experiments, moves to experiments/benchmarks. Python packaging in the next commits builds on this layout. The commit only renames and updates paths: the workspace member and the crate's path dependency, the artifact directories the benchmark tests build from CARGO_MANIFEST_DIR, the scripts that counted parent directories to find the repository, .gitignore, the README, the docs and the infra scripts. rustfmt joins three Rust lines whose strings got shorter. The crate keeps its name and binaries, and cargo package still lists only the library. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
A pattern without a slash matches at any depth, so "LICENSE", "README.md", "Cargo.toml" and "CHANGELOG.md" also packaged the matching files of every wheel in a root .venv, which uv sync creates. The patterns with a slash were already relative to the root; all of them now start with one, so the list reads the same way throughout. cargo package lists the same 79 files with or without a .venv: the library's sources and tests, the two loading docs, the manifests, README, CHANGELOG and LICENSE. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
…age. The data preparation lived in four scripts around a 1,558-line utils.py, 671 lines of which no script reached. It is now one package, experiments/treewalker_exp, built with uv_build and run as `uv run treewalker-exp`. It is a repository tool rather than a wheel: it finds the checkout by walking up to experiments/grids.toml, or takes --repo, so it also runs from outside the repository root. - grids.toml holds the grids; the package resolves them into models and cells. The factorial suite covers what `prepare.py --grid all` built, the ablation suite its anchors, and scenario-v1 the released credit cells (whatif-v1, in their released layout). - formats has the one reader and writer of each artifact file, so prepare_chunked.py and audit_f32.py now read test_data.bin through it. - paths keeps every path; pooch fetches SUPPORT, FLCHAIN and credit into experiments/data/raw/ against pinned SHA-256s, instead of a temporary directory with no checks. - compile.py runs lleaves as a subprocess under its own script lock, compile.py.lock, with a request that carries the released recipe: fblocksize 34, llc -O3, fp-contract on, the output path and the worker count. Its own defaults (O2, fp-contract fast) used to differ from the in-process call. LLVM tools come from the request or PATH. - Each compiled baseline records a compile identity next to the library: the model's hash, the settings, the compiler's resolved path, version and target triple, the host CPU (and for lleaves the CPU that llc's -mcpu=native resolves to), and tl2cgen's version or compile.py and its lock. A library is reused only when that identity and its hash match, so one built by another compiler or on another machine is rebuilt. llvmlite stays below 0.48, the last LLVM 20 release, because the VMs' llc is LLVM 20. The unused PGO, BOLT and opt paths, with their Homebrew and /usr/local paths, are gone. - tl2cgen moves into a locked baselines dependency group, and lleaves out of the project lock. The aarch64 <cstdint> compiler wrapper is a tracked script, infra/scripts/cxx-cstdint, instead of a file written to /usr/local/bin at boot. - ofat and the fixed and geometric Expedia groupings are not carried over. prepare_chunked.py stays until the runner redesign replaces it. The VM startup script and the README call the new commands. Every uv run there passes --group baselines, because uv run removes packages outside the groups it syncs. Terraform's git_ref loses its neurips2026 default, since that tag predates everything the startup script now runs. The project and compile.py declare PyPI as a uv index. A project index takes priority over indexes in a user's uv configuration, so uv sync --locked, uv lock and uv lock --check leave the locks unchanged on a machine whose configuration adds another index. prepare produces byte-identical artifacts to the base commit's scripts (on the same library versions) for SUPPORT nt500_md8_h16, Expedia nt500_md8 and the whatif-v1 credit cell k4_G16: all 30 files match a reference that the old scripts rebuilt in a clean directory, twice, with identical results. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
…ouches for.
Prep used to reuse any file that existed, so a rerun could mix old models
with new data or library versions, and nothing recorded what a cell was.
Each model now writes model.json and each cell cell.json. A model is keyed
by dataset, source hash, split, training parameters, library versions and
the training data's hash; a cell by its model, generator, parameters and
data hashes. Prep reuses one only when its key matches and every file has
its recorded hash, and rebuilds anything else. A retrained model drops its
compiled baselines, which the released runner would otherwise load.
compile-baselines records each library's request, effective settings and
hash in the model's cells.
`treewalker-exp manifest`, also run by prepare, resolves a suite into
artifacts/manifests/<suite>.json: every model, workload and cell with its
key and status, so the runner needs no grid logic. Released files keep
their places, so today's sweep_bench still reads the panel and session
cells.
New workloads, about 1,100 factorial cells in all:
- FLCHAIN horizons 256, 512 and 1,024 at T in {500, 2000} and L in {8, 16}.
- ranking-sessions-v2: each test session's candidates in one seeded
random order. The file is sorted by prop_id within each search, and the
Expedia LightGBM model at T=500, L=8 splits on prop_id in 10.0% of its
86,753 splits, so rows that a prop_id split sends the same way were
contiguous, so the rows reaching a leaf formed fewer, longer runs, and
exact sums update once per run. Requests do not arrive sorted by property. Measured on the M4 Pro (2026-10-04), with
production predict_groups over the same 10,000 test sessions: the seeded
order takes 1.3% longer than the sorted one with LightGBM, in each of 8
alternating rounds, and a median 1.4% with XGBoost, where 3 of 8 rounds
went the other way. Outputs are bit-identical once rows are mapped back.
- ranking-cohort-v2: the test sessions with at least 32 candidates. The
size-n subset (4, 8, 16, 32) is each session's first n candidates in that
seeded order, so the subsets nest and keep one order across sizes, because
row order decides how exact sums group rows into runs.
- panel-v2 declares interaction_x0_t increasing when x0 >= 0: x0 is age in
SUPPORT and FLCHAIN, fixed within a patient, while the time fraction
increases.
- whatif-v2 on all four datasets: the model's most-split eligible raw
covariates, values drawn jointly from one training entity per scenario,
and draws coupled across G and k. Survival donors are patients, not
expanded rows, which favor patients who live longer; the time step is
one seeded draw per entity, and interaction_x0_t is recomputed per row.
Each cell records the requested, selected and realized k. It is labelled
a targeted stress test; credit keeps its frozen monetary pool.
Expedia's prop_country_id, the property's country, is now declared
varying. It is the same for every candidate of 99.44% of searches, but
2,236 of the 399,344 searches with 2-128 candidates span countries, and the
session filter used to drop them so that a constant declaration held. Only
the ten search-level features are filtered on now, so the sampled sessions
change, and with them the Expedia models.
Prep checks each cell's contracts before writing its references: constant
features equal within every group, declared monotonic features monotonic
over their non-missing values. A violation fails the cell. Each model
records TreeWalker's exact import limits, 32,767 nodes per tree and a
32-million-node pool, and a model over them is marked unsupported. Each
cell records an upper bound on its varying predicates (categorical splits
are not deduplicated) and only warns above the parser's 65,535; the
runner's preflight, which loads the model, decides.
The Treelite JSON dump is written only for the models that the Rust tests
and the byte-identical check read, listed in grids.toml, and never above
the JSON loader's 64 MiB: nothing else reads it, since the runner and
tl2cgen load the binary and QuickScorer the native text. The pilot's
LightGBM model (FLCHAIN horizon 1,024, T=2000, L=16) would have written
2.53 GB. The dump is derived from the binary and tracked apart from the
model's files, so a change to the list exports or deletes it on the next
run and never retrains.
PREP_POLICY versions the rules behind every derived record: a model's
limits and JSON export, a cell's status, contracts and warnings. Cell keys
include it and model.json records it, so changing the rules re-evaluates
every model and cell on resume. A model is not retrained, and a cell whose
data and model are unchanged keeps its references; only its status and
records are recomputed.
Tests cover the generators' invariants (nesting, the shared cohort,
the session order, coupled draws, the recomputed interaction and its
monotonic declaration, realized k, the feature roles, the contracts), the resume rules (a matching manifest is reused; changed data,
library versions, models or file hashes rebuild), the limit boundaries,
predicate deduplication, the JSON allowlist, and resuming from manifests
in the two earlier formats, stored as fixtures: a cell that the first
version marked unsupported by an off-by-one predicate rule becomes ready
without --force, retraining or new references, and a stale JSON dump over
the limit is deleted. The byte-identical check passes at the previous
commit, where the 30 released files match; this one changes the panel
configurations and the Expedia sessions and models on purpose.
Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
Python 3.14 evaluates annotations lazily (PEP 649), so the import does nothing in the seven scripts that kept it. Each script still runs: paper_numbers.py reports 91 checks with no mismatches, and plot.py, gen_heatmap_tex.py, decomposition_validation.py, summarize_scenario.py, audit_f32.py and prepare_chunked.py run as before. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
The lower bounds move to lightgbm 4.7.0, xgboost 3.4.1, treelite 4.7.2, polars 1.44.2, numpy 2.5.3, pyarrow 25.0.1, plotnine 0.15.8, requests 2.34.2 and h5py 3.16.0, and uv.lock is relocked against PyPI, which also takes scipy to 1.18.1. This comes after the byte-identical check because newer libraries can train different models. - XGBoost and Treelite move together: Treelite 4.7.0 cannot import DART models from XGBoost 3.3 or later (fixed in 4.7.1). XGBoost stays at 3.4.1, the latest on PyPI. - On Linux the dependency is xgboost-cpu: the xgboost wheel there is a 57.6 MB CUDA build that pulls in NCCL, while xgboost-cpu is 5.8 MB and has no CUDA dependency. It has no macOS wheels, so macOS keeps xgboost. - Treelite 4.7.1 fixed libtreelite's rpath, so Treelite now loads libomp from Homebrew without DYLD_LIBRARY_PATH; the last instruction to set it goes. On SUPPORT nt500_md8_h16, Expedia nt500_md8 and the credit cells, data, LightGBM trees, Treelite JSON and references are unchanged. The XGBoost models change (XGBoost 3.3 rewrote the hist quantile sketch). The LightGBM text gains one parameter line, [gpu_device_id_list: ], and the Treelite binary's header records 4.7.2. The artifact correctness suite passes on these models: within 1.3e-15 for LightGBM and 6.4e-7 for XGBoost. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
The fixture generator pinned only its direct dependencies, on the command line. It now declares them as inline script metadata (Treelite 4.7.2, NumPy 2.5.3, scikit-learn 1.9.1, SciPy 1.18.1) with a committed script lock, generate.py.lock, and runs with uv run --locked --script. Like the project, the script declares PyPI as its uv index. Treelite 4.7.2 loads libomp without DYLD_LIBRARY_PATH, so that instruction goes, and the manifest's oracle comes from the installed Treelite. Every fixture and reference is regenerated. Only the binary exports change, each in one byte: the Treelite patch version in the header, 0 to 2. The JSON dumps, the GTIL and sklearn references and the inputs are byte-identical, and manifest.json records the new versions with Treelite GTIL 4.7.2 as the oracle. tests/import.rs passes unchanged. docs/treelite-loading.md moves its four 4.7.0 references to 4.7.2. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
62ca234 to
17ef4d9
Compare
rnewman
left a comment
There was a problem hiding this comment.
-
[P1] The dependency upgrade breaks VM startup.
pyproject.toml:18installs XGBoost 3.4.1, butstartup.shstill builds 3.2.0 and copies its library over the installed package. XGBoost rejects this version mismatch during import, so both VMs stop before preparation. Align the native build versions with the lockfile; LightGBM has the same discrepancy. -
[P2] Execution manifests mark stale cells ready.
manifest.py:47trusts existing cell documents without validating their files. I prepared both frameworks, changed the split seed, and rebuilt only LightGBM. The shared test data changed, invalidating XGBoost’s recorded hash, but the execution manifest still reported both cells ready. Validate artifacts before carrying forward ready. -
[P2] Workload filtering silently drops overlapping cells. grids.py:133 retains only the first workload name when deduplicating cells. Consequently,
prepare --suite factorial --workload whatif-credit-full --dry-runselects 14 cells instead of 32: the other 18 belong towhatif-credit-coreafter deduplication. Preserve workload memberships or filter before deduplicating.
The dependency upgrade moved uv.lock to LightGBM 4.7.0 and XGBoost 3.4.1, but the startup script still built 4.6.0 and 3.2.0 from source and copied them over the installed packages' libraries. XGBoost refuses a library of another version at import, so both VMs stopped before preparation. The native builds now use the locked versions. Found by Richard Newman in review. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
The execution manifest carried every cell's recorded status forward without checking its files. Cells share files, such as a model's test data: with both frameworks prepared, changing the split seed and rebuilding only LightGBM rewrote the shared test data, so XGBoost's recorded hash no longer matched, yet the manifest still reported both cells ready. A ready cell whose files no longer have their recorded hashes is now `stale`; each shared file is hashed once per manifest. Found by Richard Newman in review.
…loads share. A suite keeps one cell per ID, named for the first workload that lists it, so `--workload` missed shared cells: `prepare --suite factorial --workload whatif-credit-full --dry-run` selected 14 cells instead of 32, because the other 18 also belong to whatif-credit-core. A suite now records every workload naming each cell, and the filter matches any of them. Each cell keeps the first workload's name, so no cell's identity changes: the factorial still resolves to the same 600 models and 1,094 cells. Found by Richard Newman in review. Co-authored-by: AI (Pi/Claude Opus 5.5) <noreply@pi.dev>
The experiment scripts become one package,
treewalker-exp, and data preparation is rebuilt around manifests, so a rerun can no longer mix old models with new data. The first commit movespaper/experiments/toexperiments/andbenchmarks/toexperiments/benchmarks/, renames only.The package reproduces the released preparation byte for byte. Today's scripts, run twice from a clean checkout of #9, produced the same 30 files both times; the package's output matches them for a SUPPORT panel, an Expedia session cell and a credit what-if cell, on the old library versions. The check is pinned at the refactor commit, because the workloads commit after it changes some of those files on purpose.
model.jsonand every cellcell.json. A model is keyed by dataset, split, training parameters, library versions and the training data's hash; a cell by its model, generator, parameters and data. Prep reuses one only when its key matches and every file has its recorded hash.PREP_POLICYversions the status rules, so changing them re-evaluates old cells on resume without retraining. A compiled baseline is reused only for the same model, settings, compiler path and version, target, host CPU, and tl2cgen or llvmlite version.interaction_x0_tincreasing, since x0 is age.prop_country_idis declared varying: it is constant in 99.44% of searches, and the old prep dropped the 2,236 cross-border ones so that a constant declaration held. Candidates are now in a seeded random order per session. The file was sorted byprop_id, which the T=500, L=8 LightGBM model splits on in 10.0% of its 86,753 splits, so TreeWalker's exact sums saw fewer, longer runs: on the M4 Pro (2026-10-04) the sorted order was 1.3% faster with LightGBM in each of 8 alternating rounds, and a median 1.4% with XGBoost.xgboost-cpuon Linux, which drops CUDA and NCCL), treelite 4.7.2 and the rest at their current releases. tl2cgen moves intouv.lock, andcompile.pyruns from its own committed script lock with today's lleaves settings. PyPI is the project's first index, so a Shopify user config with the package proxy no longer rewrites the locks. Treelite 4.7.2 fixed its macOS libomp rpath, so theDYLD_LIBRARY_PATHworkaround is gone.tests/import.rspasses unchanged.The most expensive cell, FLCHAIN horizon 1,024 at T=2000, L=16, trained in 948 s with 14.3 GB peak memory on the M4 Pro (2026-10-04), for 8.2 million nodes, within every parser limit. Its Treelite JSON would have been 2.53 GB, which no loader accepts; JSON is now written only for the models the tests read. On that model
correctness.rsmisses its 1e-13 bound against GTIL at 1.21e-13. TreeWalker's own f64 full walk misses by more (2.17e-13), so this is GTIL's sequential rounding over 2,000 deep trees. The next PR's exactfsumoracle replaces that comparison.Infra:
git_refno longer defaults toneurips2026, which lacks this package, and startup uses the new commands.prepare_chunked.pystays until the runner rewrite removes its callers. The root crate's include patterns are anchored, socargo packageno longer picks up 26 license and README files from a root.venv.