On Tue, 29 Sep 2026 05:40:26 GMT, Jatin Bhateja <[email protected]> wrote:
>> Hi All, >> >> x86 AVX-512 targets provides packed shifts for 16, 32, and 64-bit lanes, but >> not for bytes. C2 lowers LShiftVB, RShiftVB, and URShiftVB by widening bytes >> to words, shifting, masking, and packing. That sequence is the byte-shift >> bottleneck. >> >> GFNI **gf2p8affineqb** instruction applies one 8×8 bit matrix to every byte. >> Logical left, logical right, and arithmetic right shifts by 0..7 are fixed >> matrices derived from the identity 0x0102040810204080. >> >> Patch uses this lowering for 128-, 256-, and 512-bit vectors when GFNI is >> available and keeps the existing widen/shift/pack path when GFNI is absent. >> >> Following are the performance numbers of benchmark included with the patch. >> >> System: AMD EPYC 9755 128-Core Processor (Turin) Fixed frequency 2.5Ghz >> >> Baseline:- >> ---------- >> Benchmark (shift) Mode Cnt Score Error Units >> ByteShiftBenchmark.ashr128 3 thrpt 2 24787.795 ops/us >> ByteShiftBenchmark.ashr256 3 thrpt 2 29184.019 ops/us >> ByteShiftBenchmark.ashr512 3 thrpt 2 42397.653 ops/us >> ByteShiftBenchmark.lshl128 3 thrpt 2 25241.434 ops/us >> ByteShiftBenchmark.lshl256 3 thrpt 2 28274.012 ops/us >> ByteShiftBenchmark.lshl512 3 thrpt 2 47331.675 ops/us >> ByteShiftBenchmark.lshr128 3 thrpt 2 25197.166 ops/us >> ByteShiftBenchmark.lshr256 3 thrpt 2 27587.841 ops/us >> ByteShiftBenchmark.lshr512 3 thrpt 2 38011.674 ops/us >> >> Withopt:- >> --------- >> Benchmark (shift) Mode Cnt Score Error Units >> ByteShiftBenchmark.ashr128 3 thrpt 2 30674.163 ops/us >> ByteShiftBenchmark.ashr256 3 thrpt 2 43275.297 ops/us >> ByteShiftBenchmark.ashr512 3 thrpt 2 72744.344 ops/us >> ByteShiftBenchmark.lshl128 3 thrpt 2 28799.612 ops/us >> ByteShiftBenchmark.lshl256 3 thrpt 2 40797.362 ops/us >> ByteShiftBenchmark.lshl512 3 thrpt 2 82156.329 ops/us >> ByteShiftBenchmark.lshr128 3 thrpt 2 30325.171 ops/us >> ByteShiftBenchmark.lshr256 3 thrpt 2 42775.756 ops/us >> ByteShiftBenchmark.lshr512 3 thrpt 2 65255.800 ops/us >> >> >> >> Kindly review and share your feedback. >> >> Best Regards, >> Jatin >> >> --------- >> - [x] I confirm that I make th... > > Jatin Bhateja has refreshed the contents of this pull request, and previous > commits have been removed. The incremental views will show differences > compared to the previous content of the PR. The pull request contains two new > commits since the last revision: > > - Rename vector_shift_gfni to vector_byte_shift_gfni > - 8393068: Optimize uniform vector byte shift operations using x86 GFNI > instruction src/hotspot/cpu/x86/c2_MacroAssembler_x86.cpp line 1372: > 1370: } > 1371: > 1372: void C2_MacroAssembler::vshiftb_gfni(int opcode, XMMRegister dst, > XMMRegister src, XMMRegister shift, Please add mathematical proof to this implementation. ------------- PR Review Comment: https://git.openjdk.org/jdk/pull/33113#discussion_r4130371273
