SamJSui opened a new pull request, #20158: URL: https://github.com/apache/tvm/pull/20158
DLight's GPU GEMV rule assumes that the final two loops around a local cache read are both owned by the cache stage. Rank-one vector inputs create only one cache loop, causing the rule to include the shared placement loop and attempt to fuse an imperfect loop nest. This change records the cache loop count before `compute_at` and recovers that innermost suffix afterward. A single cache loop is used directly, while the existing final-two-loop fusion is preserved for higher-rank inputs. A focused regression verifies that rank-one GEMV scheduling completes and that the local vector load remains vectorized. Fixes #20118 Testing: - `python -m pytest tests/python/s_tir/dlight/test_gpu_gemv.py -q` — 14 passed - `python -m pytest tests/python/relax/test_pipeline.py -q` — 24 passed - Exact issue reproducer — passed - RTX 4070 Ti SUPER (SM89): FP32/FP16 widths 2, 3, 32, 33, and 128 — NumPy parity - `pre-commit run --all-files` — passed -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
