spectrometerHBH opened a new pull request, #20103:
URL: https://github.com/apache/tvm/pull/20103

   ## Motivation and context
   
   `tirx.ptx.*` grew as a hand-written intrinsic per PTX instruction: every new
   instruction meant a Python wrapper, a codegen entry, and a hand-rolled `asm
   volatile` template, with the modifier grammar encoded implicitly in keyword
   arguments. Adding a modifier meant editing three places, and nothing checked 
the
   result against the ISA's own grammar.
   
   This replaces that surface with a table-driven dialect. Instruction shape 
lives
   in one table (mnemonic, modifier slots with their legal tokens, operand 
roles and
   dtypes); the surface, the codegen, and the type stubs are all derived from 
it.
   `T.ptx.fence.proxy.async_.shared__cta()` and `T.ptx["cp.async.bulk.tensor.2d.
   shared::cluster.global.mbarrier::complete_tx::bytes"](...)` are the same
   instruction reached two ways, and an illegal modifier is rejected at parse 
time
   with the open slots listed.
   
   ## Changes
   
   - **Table-driven PTX dialect** (`python/tvm/backend/cuda/ptx_dialect/`): 174
     instruction entries with modifier slots, operand roles/dtypes, and 
per-entry
     checks; a renderer that builds the `asm` template from the resolved chain; 
and
     generators for the coverage report and the `T.ptx` type stubs
     (`python/tvm/script/tirx.pyi`). `tirx.ptx.*` and its per-instruction
     intrinsics are retired; all in-repo call sites move to the new surface.
   - **Low-level CUDA PTX intrinsics** extended to cover the instructions the
     dialect needed (cvt/mma/ld-st families).
   - **CUDA lowering for recurrent KDA**: XOR-based swizzle address emission
     replaces the additive signed-strides path, `ldmatrix`/`stmatrix` 
destination
     registers are ordered to match the fragment word order, and register-copy 
base
     offsets are hoisted out of the serial loop.
   - **NVSHMEM objects** get CUDA's `cccl` include directory, fixing their 
build.
   - **Registry correctness test** reads the pinned bench sweep through 
whichever
     workload layout the sibling `tirx-kernels` checkout provides.
   
   ## Testing
   
   - Full `tests/python/tirx/` suite on B200 (sm_100a): 2679 passed, 73 skipped,
     3 xpassed, 0 failed, with both `tirx_kernels` import gates green.
   - `PTX_CERT=1` assembles the generated forms through `ptxas` at each entry's
     own ISA floor; `test_ptx_registration` checks every table entry is 
registered
     in a `USE_CUDA=OFF` build and `test_ptx_stub_up_to_date` byte-compares the
     checked-in stub against the generator.
   - Sphinx docs precheck locally: no new warnings from the `docs/tirx` pages.
   - `pre-commit run --all-files` clean at the pinned hook versions.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to