Skip to content

build(rocm): vendor the ROCm backend into mlxcelverse and add the rocm cargo feature #1802

Description

@inureyes

Part of #1801. Phase 0. Gates every other sub-issue.

Context

The feasibility spike in #1801 built the ROCm backend from NripeshN/mlx@75915908 on top of mlxcel's pin ml-explore/mlx@81ba1c6a, fixed six API-drift breaks, and validated it on gfx1151 (42/42 op checks, mlx-lm decode at fork parity). The result is a 121-file overlay: the fork's mlx/backend/rocm/ directory (106 files, already including the drift fixes and the mxfp4/mxfp8 scale-type dispatch fix) and 15 MLX core files merged at 81ba1c6a:

CMakeLists.txt
mlx/CMakeLists.txt
mlx/backend/common/buffer_cache.h
mlx/backend/common/compiled.cpp
mlx/backend/common/compiled.h
mlx/backend/gpu/primitives.cpp
mlx/compile.cpp
mlx/device.cpp
mlx/fast.cpp
mlx/fast.h
mlx/fast_primitives.h
mlx/io/safetensors.cpp
mlx/ops.cpp
mlx/primitives.cpp
mlx/stream.cpp

Today src/lib/mlx-cpp/CMakeLists.txt fetches MLX at the pin (line 122), applies whole-file overlays in mlx_apply_source_overlays (lines 18-82; patches-cuda/ only when MLX_BUILD_CUDA, lines 70-81), then add_subdirectory. src/lib/mlxcel-core/build.rs picks the backend by target_os plus the cuda feature, links CUDA in link_cuda() (line 492) and resolves architectures in detect_cuda_arch() (line 444). Nothing knows about ROCm.

Scope

Land the overlay and the build glue so that cargo build --features rocm produces a working GPU binary on a Linux AMD host. Nothing else in the runtime changes here; kernel routing, memory policy and quantization policy are Phase 1 sub-issues.

Implementation plan

  1. Overlay tree. The overlay is the first new member of mlxcelverse; keep it at src/lib/mlx-cpp/patches-rocm/ next to patches-cuda/ so the later reorganization (refactor(mlx-cpp): organize mlxcel's MLX-side layer as mlxcelverse (backend overlays and mlxcel kernels) #1816) moves all backends together. Add src/lib/mlx-cpp/patches-rocm/ mirroring the MLX tree with the 121 files. Add patches-rocm/UPSTREAM (source repository, branch, commit 75915908dfe5028335d318b10340313744fd3a8d, upstream base it was retargeted to, license MIT) and a short patches-rocm/README.md that states what the directory is, where it came from, and that files are whole-file overlays copied only for ROCm builds.
  2. CMake. In mlx_apply_source_overlays, copy patches-rocm/** (recursive, configure_file COPYONLY) only when MLX_BUILD_ROCM. Fail configure with a clear message if MLX_BUILD_ROCM and MLX_BUILD_CUDA are both on. Forward MLX_BUILD_ROCM, CMAKE_HIP_ARCHITECTURES and MLX_ROCM_ARCHITECTURES to the MLX subproject.
  3. Cargo features. Add rocm = [] to src/lib/mlxcel-core/Cargo.toml and rocm = ["mlxcel-core/rocm"] to the root Cargo.toml. Enabling rocm together with cuda or metal is a build error.
  4. build_mlx. When rocm is on (Linux only), define MLX_BUILD_ROCM=ON, keep MLX_BUILD_CUDA/METAL off, and pass the resolved HIP architectures.
  5. detect_rocm_arch(). Mirror resolve_cuda_architectures()/detect_cuda_arch(): honor MLX_ROCM_ARCHITECTURES if set, else parse rocminfo (Name: gfxNNNN of the GPU agents), else fail with a message telling the user to set the variable. Embed the compiled list in the binary the same way the CUDA arch list is, so a mismatch can be diagnosed at startup.
  6. link_rocm(). Resolve ROCM_PATH (default /opt/rocm). Link the static kernel library produced by the ROCm backend plus amdhip64, rocblas, hiprand, hiprtc and hipblaslt (the set the backend's own CMake links). Emit -Wl,-rpath,$ROCM_PATH/lib: the spike host does not register ROCm in the loader path, and the binary must run without LD_LIBRARY_PATH.
  7. Rebuild triggers. Add cargo:rerun-if-changed=../mlx-cpp/patches-rocm next to lines 241-242, and rerun-if-env-changed for ROCM_PATH and MLX_ROCM_ARCHITECTURES.
  8. Build cache. Confirm that a rocm build and a default build use different OUT_DIRs so the patched _deps/mlx-src is never shared. If they can collide, include the backend in the cache marker checked by purge_stale_mlx_cache (line 288).
  9. Repository gates. Run python3 scripts/ci/check_cross_repo_refs.py and rewrite any bare #NNN references in vendored comments to NripeshN/mlx#NNN or ml-explore/mlx#NNN. Run python3 scripts/insert_apache_header.py --check and confirm vendored files are skipped (all vendored .cpp/.h files carry an Apple copyright line; .hip/.hpp are outside the gate). Run make verify-kernel-dtype-keys if any vendored file contains cuda_kernel(. Add a NOTICE entry for the ROCm backend (MIT, NripeshN/mlx contributors).

Acceptance criteria

  • cargo build --release --features rocm succeeds on a Linux gfx1151 host and produces mlxcel and mlxcel-server.
  • ./target/release/mlxcel generate -m models/mlx/Qwen3-0.6B-4bit -p "Hello" -n 50 --temp 0 runs on the GPU (not the CPU fallback) and produces coherent text, with LD_LIBRARY_PATH unset.
  • A second cargo build --release --features rocm after touching a bridge .cpp reconfigures cleanly (the overlay copy is idempotent) and does not refetch MLX.
  • Editing a file under patches-rocm/ triggers a rebuild.
  • --features rocm,cuda and --features rocm,metal fail with a clear error.
  • Default, metal and cuda builds compile no file from patches-rocm/ (verify from the configure log) and their existing gates stay green.
  • check_cross_repo_refs.py and insert_apache_header.py --check pass; NOTICE and patches-rocm/UPSTREAM are present.

Validation

cargo build --release --features rocm
env -u LD_LIBRARY_PATH ./target/release/mlxcel generate -m models/mlx/Qwen3-0.6B-4bit -p "Hello" -n 50 --temp 0
readelf -d target/release/mlxcel | grep -E 'RPATH|RUNPATH|NEEDED'
touch src/lib/mlxcel-core/cpp/mlx_cxx_bridge.cpp && cargo build --release --features rocm
cargo build --release --features rocm,cuda   # must fail

Notes

  • The spike needed libopenblas-dev liblapack-dev liblapacke-dev on the host (already an mlxcel Linux prerequisite).
  • mlxcel links libmlx.a statically, so missing eval_gpu symbols fail at link time here, whereas the spike's shared Python build only failed at import. nm -D --undefined-only libmlx.so | c++filt | grep mlx::core was how the spike found them.
  • Kernel routing (refactor(core): route custom kernels by GPU backend kind instead of treating every non-Metal GPU as CUDA #1803) is not in scope. Qwen3-0.6B works without it because the affected fused paths have graph fallbacks; MoE and BitNet may not.

References

  • Overlay mechanism: src/lib/mlx-cpp/CMakeLists.txt:18-82, pin :122
  • Build glue: src/lib/mlxcel-core/build.rs (rerun-if-changed 240-242, purge_stale_mlx_cache 288, build_mlx 343, detect_cuda_arch 444, link_cuda 492)
  • Features: Cargo.toml:69-73, src/lib/mlxcel-core/Cargo.toml:12-16
  • ROCm backend link set: mlx/backend/rocm/CMakeLists.txt in the fork (target_link_libraries(mlx ...))

Activity

  1. added
    type:enhancementNew features, capabilities, or significant additions
    area:coremlxcel-core: MLX FFI, primitives, KV cache, layers
    platform:linuxLinux (CUDA / packaging) specific
    on Sep 11, 2026
  2. changed the title [-]build(rocm): vendor the xcelverse overlay and add the rocm cargo feature[/-] [+]build(rocm): vendor the ROCm backend into mlxcelverse and add the rocm cargo feature[/+] on Sep 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:coremlxcel-core: MLX FFI, primitives, KV cache, layersplatform:linuxLinux (CUDA / packaging) specificpriority:highHigh prioritystatus:doneCompletedtype:enhancementNew features, capabilities, or significant additions

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions