You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
build(rocm): vendor the ROCm backend into mlxcelverse and add the rocm cargo feature #1802
Part of #1801. Phase 0. Gates every other sub-issue.
Context
The feasibility spike in #1801 built the ROCm backend from NripeshN/mlx@75915908 on top of mlxcel's pin ml-explore/mlx@81ba1c6a, fixed six API-drift breaks, and validated it on gfx1151 (42/42 op checks, mlx-lm decode at fork parity). The result is a 121-file overlay: the fork's mlx/backend/rocm/ directory (106 files, already including the drift fixes and the mxfp4/mxfp8 scale-type dispatch fix) and 15 MLX core files merged at 81ba1c6a:
Today src/lib/mlx-cpp/CMakeLists.txt fetches MLX at the pin (line 122), applies whole-file overlays in mlx_apply_source_overlays (lines 18-82; patches-cuda/ only when MLX_BUILD_CUDA, lines 70-81), then add_subdirectory. src/lib/mlxcel-core/build.rs picks the backend by target_os plus the cuda feature, links CUDA in link_cuda() (line 492) and resolves architectures in detect_cuda_arch() (line 444). Nothing knows about ROCm.
Scope
Land the overlay and the build glue so that cargo build --features rocm produces a working GPU binary on a Linux AMD host. Nothing else in the runtime changes here; kernel routing, memory policy and quantization policy are Phase 1 sub-issues.
Implementation plan
Overlay tree. The overlay is the first new member of mlxcelverse; keep it at src/lib/mlx-cpp/patches-rocm/ next to patches-cuda/ so the later reorganization (refactor(mlx-cpp): organize mlxcel's MLX-side layer as mlxcelverse (backend overlays and mlxcel kernels) #1816) moves all backends together. Add src/lib/mlx-cpp/patches-rocm/ mirroring the MLX tree with the 121 files. Add patches-rocm/UPSTREAM (source repository, branch, commit 75915908dfe5028335d318b10340313744fd3a8d, upstream base it was retargeted to, license MIT) and a short patches-rocm/README.md that states what the directory is, where it came from, and that files are whole-file overlays copied only for ROCm builds.
CMake. In mlx_apply_source_overlays, copy patches-rocm/** (recursive, configure_file COPYONLY) only when MLX_BUILD_ROCM. Fail configure with a clear message if MLX_BUILD_ROCM and MLX_BUILD_CUDA are both on. Forward MLX_BUILD_ROCM, CMAKE_HIP_ARCHITECTURES and MLX_ROCM_ARCHITECTURES to the MLX subproject.
Cargo features. Add rocm = [] to src/lib/mlxcel-core/Cargo.toml and rocm = ["mlxcel-core/rocm"] to the root Cargo.toml. Enabling rocm together with cuda or metal is a build error.
build_mlx. When rocm is on (Linux only), define MLX_BUILD_ROCM=ON, keep MLX_BUILD_CUDA/METAL off, and pass the resolved HIP architectures.
detect_rocm_arch(). Mirror resolve_cuda_architectures()/detect_cuda_arch(): honor MLX_ROCM_ARCHITECTURES if set, else parse rocminfo (Name: gfxNNNN of the GPU agents), else fail with a message telling the user to set the variable. Embed the compiled list in the binary the same way the CUDA arch list is, so a mismatch can be diagnosed at startup.
link_rocm(). Resolve ROCM_PATH (default /opt/rocm). Link the static kernel library produced by the ROCm backend plus amdhip64, rocblas, hiprand, hiprtc and hipblaslt (the set the backend's own CMake links). Emit -Wl,-rpath,$ROCM_PATH/lib: the spike host does not register ROCm in the loader path, and the binary must run without LD_LIBRARY_PATH.
Rebuild triggers. Add cargo:rerun-if-changed=../mlx-cpp/patches-rocm next to lines 241-242, and rerun-if-env-changed for ROCM_PATH and MLX_ROCM_ARCHITECTURES.
Build cache. Confirm that a rocm build and a default build use different OUT_DIRs so the patched _deps/mlx-src is never shared. If they can collide, include the backend in the cache marker checked by purge_stale_mlx_cache (line 288).
Repository gates. Run python3 scripts/ci/check_cross_repo_refs.py and rewrite any bare #NNN references in vendored comments to NripeshN/mlx#NNN or ml-explore/mlx#NNN. Run python3 scripts/insert_apache_header.py --check and confirm vendored files are skipped (all vendored .cpp/.h files carry an Apple copyright line; .hip/.hpp are outside the gate). Run make verify-kernel-dtype-keys if any vendored file contains cuda_kernel(. Add a NOTICE entry for the ROCm backend (MIT, NripeshN/mlx contributors).
Acceptance criteria
cargo build --release --features rocm succeeds on a Linux gfx1151 host and produces mlxcel and mlxcel-server.
./target/release/mlxcel generate -m models/mlx/Qwen3-0.6B-4bit -p "Hello" -n 50 --temp 0 runs on the GPU (not the CPU fallback) and produces coherent text, with LD_LIBRARY_PATH unset.
A second cargo build --release --features rocm after touching a bridge .cpp reconfigures cleanly (the overlay copy is idempotent) and does not refetch MLX.
Editing a file under patches-rocm/ triggers a rebuild.
--features rocm,cuda and --features rocm,metal fail with a clear error.
Default, metal and cuda builds compile no file from patches-rocm/ (verify from the configure log) and their existing gates stay green.
check_cross_repo_refs.py and insert_apache_header.py --check pass; NOTICE and patches-rocm/UPSTREAM are present.
The spike needed libopenblas-dev liblapack-dev liblapacke-dev on the host (already an mlxcel Linux prerequisite).
mlxcel links libmlx.a statically, so missing eval_gpu symbols fail at link time here, whereas the spike's shared Python build only failed at import. nm -D --undefined-only libmlx.so | c++filt | grep mlx::core was how the spike found them.
changed the title [-]build(rocm): vendor the xcelverse overlay and add the rocm cargo feature[/-][+]build(rocm): vendor the ROCm backend into mlxcelverse and add the rocm cargo feature[/+]on Sep 11, 2026
Part of #1801. Phase 0. Gates every other sub-issue.
Context
The feasibility spike in #1801 built the ROCm backend from
NripeshN/mlx@75915908on top of mlxcel's pinml-explore/mlx@81ba1c6a, fixed six API-drift breaks, and validated it ongfx1151(42/42 op checks, mlx-lm decode at fork parity). The result is a 121-file overlay: the fork'smlx/backend/rocm/directory (106 files, already including the drift fixes and the mxfp4/mxfp8 scale-type dispatch fix) and 15 MLX core files merged at81ba1c6a:Today
src/lib/mlx-cpp/CMakeLists.txtfetches MLX at the pin (line 122), applies whole-file overlays inmlx_apply_source_overlays(lines 18-82;patches-cuda/only whenMLX_BUILD_CUDA, lines 70-81), thenadd_subdirectory.src/lib/mlxcel-core/build.rspicks the backend bytarget_osplus thecudafeature, links CUDA inlink_cuda()(line 492) and resolves architectures indetect_cuda_arch()(line 444). Nothing knows about ROCm.Scope
Land the overlay and the build glue so that
cargo build --features rocmproduces a working GPU binary on a Linux AMD host. Nothing else in the runtime changes here; kernel routing, memory policy and quantization policy are Phase 1 sub-issues.Implementation plan
src/lib/mlx-cpp/patches-rocm/next topatches-cuda/so the later reorganization (refactor(mlx-cpp): organize mlxcel's MLX-side layer as mlxcelverse (backend overlays and mlxcel kernels) #1816) moves all backends together. Addsrc/lib/mlx-cpp/patches-rocm/mirroring the MLX tree with the 121 files. Addpatches-rocm/UPSTREAM(source repository, branch, commit75915908dfe5028335d318b10340313744fd3a8d, upstream base it was retargeted to, license MIT) and a shortpatches-rocm/README.mdthat states what the directory is, where it came from, and that files are whole-file overlays copied only for ROCm builds.mlx_apply_source_overlays, copypatches-rocm/**(recursive,configure_file COPYONLY) only whenMLX_BUILD_ROCM. Fail configure with a clear message ifMLX_BUILD_ROCMandMLX_BUILD_CUDAare both on. ForwardMLX_BUILD_ROCM,CMAKE_HIP_ARCHITECTURESandMLX_ROCM_ARCHITECTURESto the MLX subproject.rocm = []tosrc/lib/mlxcel-core/Cargo.tomlandrocm = ["mlxcel-core/rocm"]to the rootCargo.toml. Enablingrocmtogether withcudaormetalis a build error.build_mlx. Whenrocmis on (Linux only), defineMLX_BUILD_ROCM=ON, keepMLX_BUILD_CUDA/METALoff, and pass the resolved HIP architectures.detect_rocm_arch(). Mirrorresolve_cuda_architectures()/detect_cuda_arch(): honorMLX_ROCM_ARCHITECTURESif set, else parserocminfo(Name: gfxNNNNof the GPU agents), else fail with a message telling the user to set the variable. Embed the compiled list in the binary the same way the CUDA arch list is, so a mismatch can be diagnosed at startup.link_rocm(). ResolveROCM_PATH(default/opt/rocm). Link the static kernel library produced by the ROCm backend plusamdhip64,rocblas,hiprand,hiprtcandhipblaslt(the set the backend's own CMake links). Emit-Wl,-rpath,$ROCM_PATH/lib: the spike host does not register ROCm in the loader path, and the binary must run withoutLD_LIBRARY_PATH.cargo:rerun-if-changed=../mlx-cpp/patches-rocmnext to lines 241-242, andrerun-if-env-changedforROCM_PATHandMLX_ROCM_ARCHITECTURES.rocmbuild and a default build use differentOUT_DIRs so the patched_deps/mlx-srcis never shared. If they can collide, include the backend in the cache marker checked bypurge_stale_mlx_cache(line 288).python3 scripts/ci/check_cross_repo_refs.pyand rewrite any bare#NNNreferences in vendored comments toNripeshN/mlx#NNNorml-explore/mlx#NNN. Runpython3 scripts/insert_apache_header.py --checkand confirm vendored files are skipped (all vendored.cpp/.hfiles carry an Apple copyright line;.hip/.hppare outside the gate). Runmake verify-kernel-dtype-keysif any vendored file containscuda_kernel(. Add aNOTICEentry for the ROCm backend (MIT, NripeshN/mlx contributors).Acceptance criteria
cargo build --release --features rocmsucceeds on a Linuxgfx1151host and producesmlxcelandmlxcel-server../target/release/mlxcel generate -m models/mlx/Qwen3-0.6B-4bit -p "Hello" -n 50 --temp 0runs on the GPU (not the CPU fallback) and produces coherent text, withLD_LIBRARY_PATHunset.cargo build --release --features rocmafter touching a bridge.cppreconfigures cleanly (the overlay copy is idempotent) and does not refetch MLX.patches-rocm/triggers a rebuild.--features rocm,cudaand--features rocm,metalfail with a clear error.metalandcudabuilds compile no file frompatches-rocm/(verify from the configure log) and their existing gates stay green.check_cross_repo_refs.pyandinsert_apache_header.py --checkpass;NOTICEandpatches-rocm/UPSTREAMare present.Validation
Notes
libopenblas-dev liblapack-dev liblapacke-devon the host (already an mlxcel Linux prerequisite).libmlx.astatically, so missingeval_gpusymbols fail at link time here, whereas the spike's shared Python build only failed at import.nm -D --undefined-only libmlx.so | c++filt | grep mlx::corewas how the spike found them.References
src/lib/mlx-cpp/CMakeLists.txt:18-82, pin:122src/lib/mlxcel-core/build.rs(rerun-if-changed240-242,purge_stale_mlx_cache288,build_mlx343,detect_cuda_arch444,link_cuda492)Cargo.toml:69-73,src/lib/mlxcel-core/Cargo.toml:12-16mlx/backend/rocm/CMakeLists.txtin the fork (target_link_libraries(mlx ...))