Skip to content

Test one kernel driving two tensors through two peer tables - #552

Merged
mawad-amd merged 13 commits into
mainfrom
muhaawad/unify-test
Oct 5, 2026
Merged

mawad-amd merged 13 commits into
mainfrom
muhaawad/unify-test

Conversation

@mawad-amd

Copy link
Copy Markdown
Collaborator

Adds a single test parameterized over providers. It allocates two
symmetric tensors, takes each one's own peer table, and runs one kernel
that reads the peer's first tensor through the first table and writes the
peer's second tensor through the second.

Technical Details

Test Plan

Test Result

Submission Checklist

mawad-amd and others added 2 commits September 22, 2026 08:32
Adds a single test parameterized over providers. It allocates two
symmetric tensors, takes each one's own peer table, and runs one kernel
that reads the peer's first tensor through the first table and writes the
peer's second tensor through the second.

What that reaches which a single-tensor test cannot: each tensor resolves
only through its own table. Iris and rocSHMEM both allocate from one
symmetric heap, so the peer delta is a constant of the heap and any table
happens to translate any pointer. Torch Symmetric Memory is symmetric
memory rather than a symmetric heap -- each tensor is its own allocation
with its own peer mapping -- so it has one translation per tensor and the
deltas differ between allocations. A kernel reusing one table for two
tensors works on the heap providers and silently mistranslates on the
per-tensor one.

Adding a provider is one entry in PROVIDERS; Torch Symmetric Memory joins
when its provider lands.

Providers are built inside the fixture rather than at module scope so an
unavailable one skips its own parameters instead of collecting zero items,
which would make pytest exit 5 and fail the distributed run.
@mawad-amd
mawad-amd marked this pull request as draft September 22, 2026 15:56
@github-actions github-actions Bot added in-progress We are working on it iris Iris project issue labels Sep 22, 2026
The device-side contract is a plain int64 table, so it does not depend on
which backend consumes it. Adds a Gluon kernel that does the same manual
translation and parameterizes the test over both, so the same two tables
are asserted to drive either dialect to the same result.

Sizes the BlockedLayout from BLOCK_SIZE rather than hardcoding [1]: the
existing Gluon tests use blocks of 32 or fewer, where one element per lane
covers the block. At 256 it does not.
The Gluon BlockedLayout has to agree with the block, and the lane count in
it is not a constant: 64 on CDNA, 32 on NVIDIA and RDNA. Query it with
triton.runtime.driver.active.get_current_target().warp_size and pass it in
as a constexpr, rather than hardcoding the CDNA value.

Asserts the block divides evenly into warps, so a bad pairing fails with a
clear message rather than as a layout compile error.
peer_ptrs[r] is the tensor's address on rank r, so subtracting our own
entry from our own pointer is subtracting a number from itself. The
byte-cast and add cancelled with it. Carried over from the heap-anchored
kernel, where the anchor names the heap rather than the allocation and the
offset is real.

The kernel no longer needs its own rank, so CUR_RANK goes too.
Every method on it forwarded unchanged. Iris.allocate_symmetric,
get_rank, get_num_ranks and barrier already match what the rocSHMEM
provider exposes, signatures included, so the context goes into PROVIDERS
unwrapped.

Needing an adapter would have meant the two interfaces had not actually
converged. Also drops a .name attribute nothing read.
@mawad-amd
mawad-amd marked this pull request as ready for review September 22, 2026 21:45
Comment thread tests/unittests/test_provider_unified.py Outdated
Comment thread tests/unittests/test_provider_unified.py Outdated
Both provider test modules called init_rocshmem_by_uniqueid in their own
module-scoped fixture, so running them in one pytest process initialised
the runtime twice. The runtime is per-process; the existing module scope
was chosen for exactly that reason and adding a second module broke the
assumption.

Moves the init to a session-scoped rocshmem_runtime fixture in
tests/unittests/conftest.py and has both modules depend on it. One init
per process however many modules ask. Still no finalize in teardown --
it would pull the runtime out from under anything still running.

Raised by @drprajap on #552.
The test allocated two symmetric tensors and never released them, so a
failing assert leaked both. Adds a fixture that records what it hands out
and releases it after the yield, which pytest runs on failure as well as
success.

Release goes through a _deallocate stub: rocSHMEM frees explicitly because
rocshmem_free is collective, Iris has no deallocate yet. When it grows
one, this is the single place to change.

Raised by @drprajap on #552.
# Conflicts:
#	tests/unittests/test_rocshmem_provider.py
@mawad-amd
mawad-amd merged commit 57cba63 into main Oct 5, 2026
51 checks passed
@mawad-amd
mawad-amd deleted the muhaawad/unify-test branch October 5, 2026 11:05
nirvedhmeshram added a commit to nirvedhmeshram/iris that referenced this pull request Oct 9, 2026
…was measured

Register the torch provider in test_provider_unified.py (ROCm#552). Its probe
and skips move into a session fixture in conftest so both modules share
them, and the two torch tests the unified test now covers -- the table
invariant and per-allocation anchoring -- are dropped.

The backend guidance overreached. "Do not call set_backend" and "forcing
any other name fails at allocation" generalised from forcing NCCL on 2.13
builds; on the 2.15 rocm10.0 nightly symmetric memory reportedly works only
with set_backend("NVSHMEM"), i.e. rocSHMEM. The provider never chose a
backend anyway, so the docs now say the choice is the caller's, keep the
one real trap (set_backend('CUDA') is rejected even though get_backend
reports it), and record the nightly report as reported, not measured.

Likewise "peer offsets are not shared between allocations" is a property
of the default IPC backend, not of torch symmetric memory; reworded as
"need not be". Fixes the README saying allocation is collective while the
docstring said it is local: only rendezvous is, on the default backend.

The module docstring is cut to the points that matter when reading the
code; the rest already lives in the README.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

in-progress We are working on it iris Iris project issue

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants