spectrometerHBH opened a new pull request, #20153: URL: https://github.com/apache/tvm/pull/20153
This PR adds `T.ptx.addr(base, byte_offset)` to the TIRx PTX dialect: a pure expression that folds a compile-time signed byte displacement into the PTX address operand, rendering as `[%N+imm]` instead of requiring a separate address computation before the instruction. - `tirx.ptx.addr` is an expression only in the outer PTX call's IR: the helper still receives the coerced base register, while the displacement becomes renderer metadata baked into the instruction text and the helper name (`_addr<slot>_p<imm>` / `_m<imm>`). - Table-level `allow_imm_offset` classification of address slots, with validation that rejects immediate offsets on operand classes that cannot take them (e.g. `tmem` addresses). - Immediate operands are now validated to be compile-time `IntImm` at CUDA codegen, with an actionable error pointing at explicitly-unrolled loops. - Displacements are range-checked to int32; zero offsets normalize to the bare form so existing helper names are untouched. Also includes a small test fix: `test_tirx_kernels_registry_correctness.py` accepts both the old and new MegaMoE kernel registry names (`deepgemm_fp8_fp4_mega_moe` / `sm100_fp8_fp4_mega_moe`), so the test works against tirx-kernels checkouts from either side of the rename. Tested with `tests/python/tirx/codegen/test_ptx_addr.py` (new, 12 cases), plus the full `tests/python/tirx/` suite on sm100. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
