jinhongyii opened a new pull request, #20155:
URL: https://github.com/apache/tvm/pull/20155

   CUDA tensor-map enum members are C++ identifiers, not preprocessor macros. 
The existing `#ifdef` checks therefore evaluated false even when CUDA 12.8+ 
`cuda.h` declared packed tensor-map data types and atomic 128-byte swizzle 
modes. This caused host validation to reject supported swizzles and to skip 
their packed/128-byte validation rules.\n\nThis change uses the existing 
`CUDA_VERSION >= 12080` availability contract consistently for every affected 
dtype and swizzle validation path. Older toolkits still compile without 
referencing the newer enum members.\n\nTesting:\n- `pre-commit run clang-format 
--files src/backend/cuda/runtime/cuda_device_api.cc`\n- CUDA 13 build with 
`USE_NVSHMEM=ON`: `cmake --build build --parallel 16`\n- Downstream TIRx GDN 
swizzle path: 10 correctness launches and 3 benchmark rounds passed


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to