Repository navigation
Conversation
block_bucketize_sparse_features(_inference) gave each row to one work-item that walked its indices in order. Batch-1 inference, as in TorchRec row-wise sharding for DLRM-v3, has a few rows with thousands of indices, so the kernel ran on a handful of work-items: 25 ms for one 16k-index row on Data Center GPU Max, against 0.12 ms on the CPU. Add a chunked path. Rows are split into chunks of one sub-group (32 indices). One kernel counts the buckets of each chunk, a column-wise cumsum over chunks gives every chunk its offset within its (bucket, row) and new_lengths, and the scatter ranks each index within its chunk and bucket with a sub-group ballot. Positions follow the input order, so the output equals the serial kernels and the CPU kernel bit for bit, for pooled rows too. The bucket arithmetic repeats the serial kernels. The serial kernels stay unchanged and remain the default for many or short rows, where they already fill the device. The chunked path is chosen for at most 8192 rows with a mean length of at least 128, the range where it was faster in a sweep on Data Center GPU Max. FBGEMM_XPU_BLOCK_BUCKETIZE_KERNEL=serial|chunked forces one path. Data Center GPU Max, one tile, sequence=True, return_bucket_mapping=True, 8 buckets: 256 / 2048 / 16384 indices in one row 0.48 / 3.3 / 25.8 ms -> 0.24 / 0.27 / 0.27 ms; three 16k rows with uneven buckets 36.6 -> 0.44 ms. Signed-off-by: mkrze <mateusz.krzemieniewski@intel.com>
Compare every output of block_bucketize_sparse_features(_inference) with FBGEMM's CPU kernels for the serial path, the chunked path and the automatic choice. Row lengths straddle the 32-index chunk (0, 1, 31, 32, 33) and reach 16384, with 1 to 64 buckets, int32/int64 lengths and indices, weights, bucketize_pos, keep_orig_idx (global and per feature), total_num_blocks, uneven buckets, raw and out-of-range ids, variable batch sizes and a seeded random sweep. populate_bucketized_permute must reproduce the chunked unbucketize_permute. Add a long-row case for both paths to the non-current-device test. Signed-off-by: mkrze <mateusz.krzemieniewski@intel.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft, in progress