jinhongyii opened a new pull request, #20155: URL: https://github.com/apache/tvm/pull/20155
CUDA tensor-map enum members are C++ identifiers, not preprocessor macros. The existing `#ifdef` checks therefore evaluated false even when CUDA 12.8+ `cuda.h` declared packed tensor-map data types and atomic 128-byte swizzle modes. This caused host validation to reject supported swizzles and to skip their packed/128-byte validation rules.\n\nThis change uses the existing `CUDA_VERSION >= 12080` availability contract consistently for every affected dtype and swizzle validation path. Older toolkits still compile without referencing the newer enum members.\n\nTesting:\n- `pre-commit run clang-format --files src/backend/cuda/runtime/cuda_device_api.cc`\n- CUDA 13 build with `USE_NVSHMEM=ON`: `cmake --build build --parallel 16`\n- Downstream TIRx GDN swizzle path: 10 correctness launches and 3 benchmark rounds passed -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
