feat(rocm): add RDNA3 and RDNA4 runtime foundation - #132
Open
zihaomu wants to merge 8 commits into
Open
Conversation
- Add hip_compat.h shim mapping CUDA runtime API to HIP equivalents - Update pinned_tensor.cpp to compile under both nvcc and hipcc - Add ROCm detection in arch.py (is_rocm, get_rocm_gfx_arch, is_gfx11xx_family) - Guard NVIDIA arch checks to return None on ROCm - Skip nvcc version check in _toolchain.py when on ROCm - Add ROCm build path in setup.py (ROCM_HOME, amdhip64, --offload-arch) - Add _hip_cflags() in kernel/utils.py for JIT compilation on ROCm - Add is_rocm() and driver_hip_version() in backend.py - Add rocm-smi fallback in __main__.py for clangd generation - Add TODO(ROCm) for NCCL->RCCL, flashinfer/sgl_kernel ROCm builds, Triton autotune RDNA3 tuning, PDL equivalent, hiprtc JIT cache - Add AMD ROCm classifier in pyproject.toml
This was referenced Aug 24, 2026
Author
|
Draft follow-ups are now available:
All four are opened as Draft PRs against |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
gfx1100-gfx1103) and RDNA4 (gfx1200,gfx1201);/opt/rocmlayouts and the modular ROCm SDK shipped by the official ROCm 7.14 PyTorch image;gcnArchName, with explicit environment overrides for cross compilation.Why this targets
maindirectlyThis PR is a standalone, current-
mainintegration. It incorporates the useful work from #23 and the RDNA3 runtime follow-up at bouclem#1, then adds the RDNA4 and ROCm 7.14 changes. It therefore does not require either external feature branch to merge first.This intentionally overlaps those open PRs so maintainers can review and merge a complete, hardware-tested ROCm foundation without being blocked by a cross-fork base chain.
Attribution
7caa62d(upstream source commit27c0977).af67560through62bb962(upstream follow-up head4e11ed0).Thank you to both contributors for establishing and validating the earlier ROCm paths.
Validation
Hardware and software:
gfx1201)2.11.0+rocm7.14.07.14.608503.7.1Results:
python -m pip install -e . --no-build-isolation --no-deps: PASS; both native extensions compiled and loaded;48 passed;python -m pip check: PASS;git diff --check: PASS.Compatibility
ROCm-specific code is guarded by HIP/ROCm detection and existing CUDA paths are retained. Physical NVIDIA regression testing is not claimed; upstream CUDA CI is requested before merge.
Follow-ups
Kept out of this foundation PR for separate review: