Suggestion Description
#551 provides capability to run Iris device kernels with memory allocated via PyTorch, with no copy, no second heap, and no kernel changes.
and we agreed on extending it to accept tensors frameworks have already rendezvoused , because users might want to manage the memory.
Overall Motivation behind #551 and this feature request is as follows:
- Some of the inference stack already uses torch.distributed.symmetric_memory. TokenSpeed's amd_rsag* path, the TritonAllReduceBackend, and the torch.ops.symm_mem.* collectives that vLLM and SGLang-style engines may call all allocate through it. Until now, using an Iris kernel meant moving those buffers into the Iris heap. This provider lets the two coexist: a framework keeps its torch allocations, and individual Iris kernels can run on them.
- Frameworks that already own their buffers can't use Iris device kernels until the follow-up lands.
- The single-map
iris.copy limitation also means it does not yet make every Iris operation work across independently allocated tensors
Operating System
No response
GPU
No response
ROCm Component
No response
Suggestion Description
#551 provides capability to run Iris device kernels with memory allocated via PyTorch, with no copy, no second heap, and no kernel changes.
and we agreed on extending it to accept tensors frameworks have already rendezvoused , because users might want to manage the memory.
Overall Motivation behind #551 and this feature request is as follows:
iris.copylimitation also means it does not yet make every Iris operation work across independently allocated tensorsOperating System
No response
GPU
No response
ROCm Component
No response