Skip to content

fbgemm-xpu: adds performance benchmarking for the XPU jagged tensor operators - #142

Draft
aagalleg wants to merge 14 commits into
intel:mainfrom
aagalleg:bench/jagged_ops
Draft

aagalleg wants to merge 14 commits into
intel:mainfrom
aagalleg:bench/jagged_ops

Conversation

@aagalleg

@aagalleg aagalleg commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Adds a standalone jagged-sweep benchmark that verifies correctness against CPU before timing, patches the upstream FBGEMM jagged benchmark to run on XPU, commits PVC and BMG baselines with notes on interpreting them, and wires smoke runs into CI. Together these provide a reproducible performance reference for detecting regressions in the SYCL jagged kernels.

Operator(s)

Benchmarked (forward and backward; previously implemented):

  • fbgemm::jagged_to_padded_dense — copies jagged values into a padded dense tensor
  • fbgemm::jagged_2d_to_dense — 2D variant of jagged→dense conversion
  • fbgemm::dense_to_jagged — gathers valid rows of a dense tensor into jagged form
  • fbgemm::jagged_dense_elementwise_add_jagged_output — adds dense to jagged values, jagged output

Changes

  • jagged-sweep benchmark in src/fbgemm_xpu/bench/jagged_sweep.py: sweeps B ∈ {32…2048} × max_len ∈ {16…1024} × {float32, bfloat16, float16}, D = 128 (DLRM v3 hstu_transducer_embedding_dim); every shape is verified against the CPU implementation before timing; emits CSV with a metadata header
  • Timing helper in src/fbgemm_xpu/bench/bench_utils.py: XPU-event timing with last-level-cache flush and a GPU lead (torch.xpu._sleep spin) so the timed window never measures host submission
  • XPU skip handling for non-ported upstream benchmark cases in src/fbgemm_xpu/bench/skips.py
  • Patch 0002-Add-XPU-support-to-fbgemm-jagged-benchmark.patch enabling the upstream jagged_tensor_benchmark.py device case on XPU
  • Committed baselines: bench/baselines/pvc-max1550/jagged_tensor.csv and bench/baselines/bmg-b60/jagged_tensor.csv, plus bench/baselines/README.md documenting methodology, byte-count formulas, and known caveats (effective vs. physical bandwidth, latency floors, BMG memory compression)
  • CI: smoke-runs jagged-sweep (correctness gate) and the patched upstream benchmark on every run; validates the patch touches only fbgemm_gpu/bench/
  • Docs: CONTRIBUTING.md benchmarking section and README.md updates
  • Tests in tests/test_jagged_sweep.py, tests/test_bench_utils.py, tests/test_bench_skips.py

Testing

pytest -rsf packages/fbgemm-xpu/tests/   # includes test_jagged_sweep, test_bench_utils, test_bench_skips
python -m fbgemm_xpu.bench.jagged_sweep --smoke --output /tmp/jagged_sweep_smoke.csv
python fbgemm_gpu/bench/jagged_tensor_benchmark.py device --batch-size 8 --max-len 16   # after patch

jagged-sweep compares every forward output and input gradient against the CPU implementation with torch.testing.assert_close before timing; the run fails and writes no CSV on any mismatch. All 576 sweep cases passed on both devices.

Baseline run-to-run variation: PVC median 0.7% (max 5.5%), BMG median 0.1% (max 4.9%).

Hardware Validated

  • Intel Data Center GPU Max 1550 (PVC), one tile (ZE_AFFINITY_MASK=0)
  • Intel Arc Pro B60 Graphics (BMG), one device, 4-CPU pod with OMP_NUM_THREADS=4

@aagalleg aagalleg changed the title Adds performance benchmarking for the XPU jagged tensor operators fbgemm-xpu: adds performance benchmarking for the XPU jagged tensor operators Oct 6, 2026
Mark every assert in the bench test modules with "# nosec B101",
matching the convention of the other test files: pytest never runs
under python -O, so the asserts are not stripped in practice.

Also annotate the subprocess usage in test_jagged_sweep.py with
"# nosec B404"/"# nosec B603". The test invokes sys.executable with
a fixed argument list and no shell, so no untrusted input reaches
the call; the suppression comments record that audit.
… nbit tests

The int nbit lookup fixtures required exactly torch 2.14.0 and
fbgemm-gpu-cpu 1.9.0, so every test using them errored on torch
2.14.1. Check only the major and minor versions, which matches the
torch~=2.14.0 pin in pyproject.toml, and report the installed version
when the check fails.
_copy_bytes_per_s copied from torch.empty, and freshly allocated
device memory is typically zero-filled. Zeros are the best case for
memory compression, so the copy can report more than the device's
real bandwidth: on the BMG runner the streaming copy measured
518 GB/s, above the part's nominal ~456 GB/s.

Fill the source with torch.rand so both copies in
test_default_flush_evicts_last_level_cache measure incompressible
traffic.
The fbgemm-xpu test step has hung on a runner with no output beyond
the test file name, so there is no way to tell which call blocked.
Pass -o faulthandler_timeout=600 to pytest: a test still running
after 10 minutes dumps every thread's stack into the job log, which
points at the blocking line.

Reporting only; the hung test is not interrupted and the job still
runs to its own timeout.
@@ -0,0 +1,594 @@
# command: python -m fbgemm_xpu.bench.jagged_sweep --output jagged_tensor.csv

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suggest to drop comments in the .csv file and have only actual data here. This way Github will actually be able to render the .csv file for us in a nice column way. It will actually become reviewable. If you need to save comments, then create a README.md in the folder and describe what each .csv file contains there.

…s limits

The fbgemm-xpu test job runs two matrix entries at once on one BMG
runner with a 12 GB GPU. Three things then went wrong, all in CI only.

The flush test compared a copy sized to fit the last-level cache against
a streaming copy 32x larger and expected them within 25%. That ratio
encodes PVC behaviour: on GDDR6 a long read+write stream runs well
below peak, while a short copy's writes are absorbed by the write-back
cache, which no flush ahead of the copy can prevent. On BMG the two
differ by 1.8x with the flush working correctly. Compare the same copy
with and without the default flush instead; the cached copy must be at
least 1.5x faster. PVC gives 4.3x, and upstream's fixed 40 MB flush
still fails at 0.78x.

The bench tests left 2.4 GiB in the caching allocator for the rest of
the session, and the large-grid tests allocate 8 GiB operands behind a
free-memory check that cannot see the other process. The driver backs
allocations past physical memory with system memory instead of
failing, so both jobs crawled rather than raised. Release the bench
operands after each test, halve the bandwidth-bound test's working set
to 1 GiB, and skip any large-operand test whose footprint exceeds half
of the device's physical memory, leaving room for a second process.

benchmark_torch_function sized its GPU lead from the slowest warm-up,
which includes one-time kernel compilation (8-112 ms against a 20-80 us
steady submission), so the lead hit its 50 ms cap, and upstream call
sites queue it 1000 times: 50 s of GPU spin per measurement, up to
200 s with retries, and a GPU the other job could not use. Size the
lead from the fastest warm-up, which the retry loop already covers if
too small, and cap the lead queued across all iterations at 1 s. The
patched upstream benchmark's smoke run drops from 130 s to 9 s with
timings unchanged. Only warn about host waits once they reach half of
the iterations, where they can move the median; a few jittery
iterations in 1000 no longer flag every measurement.
c4fe413 passed -o faulthandler_timeout=600 to pytest so that a test
still running after 10 minutes would dump every thread's stack and
show where the job had stalled. The stall was two test processes
oversubscribing the shared GPU's memory, and 00436ef removes its
causes, so the job no longer needs the diagnostic.

This reverts commit c4fe413.
test_default_flush_evicts_last_level_cache compares a half-LLC copy
timed with and without the default flush and expects the flushed copy
to be at least 1.5x slower. On the BMG runner the two came out equal at
~71 GB/s, 128 us for a 9 MB copy the part streams in ~25 us. The job
shares its GPU with the other matrix entry, and time-slicing adds the
same latency to both copies, so the cache effect the test looks for is
buried whether or not the flush works. Reproduced on PVC with a matmul
in another process: both copies at 43.6 ms, ratio 1.00.

Add test_default_flush_is_twice_last_level_cache, which spies on
Tensor.zero_ to check that the helper sizes its scratch buffer to twice
the device's last-level cache and overwrites it before every timed
iteration and not before warm-ups. This holds on any part and on a
shared GPU, and still catches a return to upstream's fixed 40 MB.

Have the bandwidth test measure a launch floor first, a one-element
copy, and skip with the measured numbers when the cached copy is not
clearly above it: either the GPU is shared or the cache is too small to
resolve, and in both cases the comparison says nothing about the flush.
On an idle PVC it still runs and passes with a 4.2x ratio.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants