SamJSui opened a new pull request, #20160: URL: https://github.com/apache/tvm/pull/20160
Fixes #20060. ## What changed - Detect reduction temporaries that are private to a single scalar producer kernel. - Assign those buffers `local` scope only when all producer loops have unit extent and no other block accesses the buffer. - Preserve `global` scope for temporaries consumed by another block or kernel. - Add a focused regression covering reduction lengths one, two, and three. ## Root cause Length-one scalar `argmin` produces paired index and value temporaries. The index is consumed by the final result block and must retain global workspace lifetime. The value temporary ends in the producer kernel, but DLight left it in global scope. CUDA source generation rejects that internal global allocation. The fix is lifetime-based rather than specific to `argmin` or to a positional output buffer. ## Validation - Base regression fails because the private value temporary remains `global`; patched regression passes. - Exact current-API CUDA reproducer: base fails during CUDA source generation; patched build passes. - DLight and adjacent Relax suites: `153 passed, 1 skipped, 1 xfailed`. - CUDA trigger matrix: 11/11 cases pass, including `argmin`, `argmax`, scalar/all-axis reductions, `float16`, `float32`, and `int32`. - Mixed-order and tie cases match NumPy. - OpenCL source generation emits a local value temporary while preserving the global index parameter. - All applicable changed-file pre-commit hooks pass. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
