jinhongyii opened a new pull request, #20100: URL: https://github.com/apache/tvm/pull/20100
NVRTC uses `--gpu-architecture` to select the frontend target. Passing the real-architecture shorthand (for example, `sm_100a`) may compile through a generic virtual target and reject architecture-family-specific inline PTX, while the nvcc path already separates `arch=compute_<suffix>` from `code=sm_<suffix>`. This change makes the NVRTC path: - use `compute_<suffix>` for auto-detected targets; - normalize an explicitly supplied `sm_*` target to the corresponding `compute_*` target. Validation: - `pre-commit run --files python/tvm/support/nvcc.py` - CUDA 13.2 / B200: compiled the full sparse FlashMLA V32 kernel for `sm_100a` through the default NVRTC path, including its direct `cvt.rn.bf16x2.e4m3x2` instruction. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
