spectrometerHBH opened a new pull request, #20153:
URL: https://github.com/apache/tvm/pull/20153

   This PR adds `T.ptx.addr(base, byte_offset)` to the TIRx PTX dialect: a pure 
expression that folds a compile-time signed byte displacement into the PTX 
address operand, rendering as `[%N+imm]` instead of requiring a separate 
address computation before the instruction.
   
   - `tirx.ptx.addr` is an expression only in the outer PTX call's IR: the 
helper still receives the coerced base register, while the displacement becomes 
renderer metadata baked into the instruction text and the helper name 
(`_addr<slot>_p<imm>` / `_m<imm>`).
   - Table-level `allow_imm_offset` classification of address slots, with 
validation that rejects immediate offsets on operand classes that cannot take 
them (e.g. `tmem` addresses).
   - Immediate operands are now validated to be compile-time `IntImm` at CUDA 
codegen, with an actionable error pointing at explicitly-unrolled loops.
   - Displacements are range-checked to int32; zero offsets normalize to the 
bare form so existing helper names are untouched.
   
   Also includes a small test fix: `test_tirx_kernels_registry_correctness.py` 
accepts both the old and new MegaMoE kernel registry names 
(`deepgemm_fp8_fp4_mega_moe` / `sm100_fp8_fp4_mega_moe`), so the test works 
against tirx-kernels checkouts from either side of the rename.
   
   Tested with `tests/python/tirx/codegen/test_ptx_addr.py` (new, 12 cases), 
plus the full `tests/python/tirx/` suite on sm100.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to