spectrometerHBH opened a new pull request, #20159: URL: https://github.com/apache/tvm/pull/20159
## Summary - add a `tirx.max_registers` PrimFunc attribute and lower it to CUDA 13 `__maxnreg__` - reject combinations of the register cap with CUDA launch-bounds attributes - enable non-portable CUDA cluster sizes when a launch requests more than eight CTAs - keep launch attributes in fixed local storage and add CUDA codegen coverage ## Motivation SM100 TIRx kernels need an explicit register cap to preserve source-level occupancy and instruction scheduling. Large reduction kernels can also require 16-CTA clusters, which the CUDA driver rejects unless the non-portable cluster-size attribute is enabled. ## Impact TIRx kernels can opt into an exact CUDA register budget while retaining the existing launch-bounds path for other specializations. CUDA launches with cluster dimensions above eight are enabled only for kernels that request them; existing launch behavior remains unchanged. ## Validation - `cmake --build build --parallel` - `python -m pytest tests/python/tirx/codegen/test_codegen_cuda.py -q` (`201 passed`) - `pre-commit run --files include/tvm/tirx/function.h src/backend/cuda/codegen/codegen_cuda.cc src/backend/cuda/runtime/cuda_module.cc src/tirx/transform/split_host_device.cc tests/python/tirx/codegen/test_codegen_cuda.py` -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
