SamJSui opened a new pull request, #20158:
URL: https://github.com/apache/tvm/pull/20158

   DLight's GPU GEMV rule assumes that the final two loops around a local cache
   read are both owned by the cache stage. Rank-one vector inputs create only 
one
   cache loop, causing the rule to include the shared placement loop and attempt
   to fuse an imperfect loop nest.
   
   This change records the cache loop count before `compute_at` and recovers 
that
   innermost suffix afterward. A single cache loop is used directly, while the
   existing final-two-loop fusion is preserved for higher-rank inputs.
   
   A focused regression verifies that rank-one GEMV scheduling completes and 
that
   the local vector load remains vectorized.
   
   Fixes #20118
   
   Testing:
   
   - `python -m pytest tests/python/s_tir/dlight/test_gpu_gemv.py -q` — 14 
passed
   - `python -m pytest tests/python/relax/test_pipeline.py -q` — 24 passed
   - Exact issue reproducer — passed
   - RTX 4070 Ti SUPER (SM89): FP32/FP16 widths 2, 3, 32, 33, and 128 — NumPy 
parity
   - `pre-commit run --all-files` — passed
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to