https://github.com/wangpc-pp created https://github.com/llvm/llvm-project/pull/215517
In `Zvdot4a8i` extension, `vdot4a*` only produces an i32 result. This PR extends partial-reduction handling so an i64 accumulator with i8 inputs (a scale-8 partial reduction) is also lowered using the dot-product instructions: perform the dot product into a fresh i32 accumulator (each lane is the sum of four i8 products and cannot overflow i32), then sign/zero-extend the i32 partial sums to i64 and accumulate. This mirrors the AArch64 sdot+sadalp idiom. - `getPartialReductionCost` accepts an i64 accumulator (reduction factor 8), costing the `vdot4a*` plus the extra i32->i64 widen and accumulate. - `setPartialReduceMLAAction` marks the i8->i64 (scalable and fixed-length) partial reductions Custom. - `lowerPARTIAL_REDUCE_MLA` lowers the i64 case via an i32 `vdot4a*` followed by a widening `partial.reduce.add`, which the generic expander turns into extract_subvector + add. The i64 case requires a wide enough VF (LMUL) to reach the scale-8 factor, matching how AArch64 only forms it under SVE. Assisted-by: TRAE CLI (Opus 4.8) _______________________________________________ llvm-branch-commits mailing list [email protected] https://lists.llvm.org/cgi-bin/mailman/listinfo/llvm-branch-commits
