On Tue, 29 Sep 2026 05:40:26 GMT, Jatin Bhateja <[email protected]> wrote:

>> Hi All,
>> 
>> x86 AVX-512 targets provides packed shifts for 16, 32, and 64-bit lanes, but 
>> not for bytes. C2 lowers LShiftVB, RShiftVB, and URShiftVB by widening bytes 
>> to words, shifting, masking, and packing. That sequence is the byte-shift 
>> bottleneck.
>> 
>> GFNI **gf2p8affineqb** instruction applies one 8×8 bit matrix to every byte. 
>> Logical left, logical right, and arithmetic right shifts by 0..7 are fixed 
>> matrices derived from the identity 0x0102040810204080.
>> 
>> Patch uses this lowering for 128-, 256-, and 512-bit vectors when GFNI is 
>> available and keeps the existing widen/shift/pack path when GFNI is absent.
>> 
>> Following are the performance numbers of benchmark included with the patch.
>> 
>> System: AMD EPYC 9755 128-Core Processor (Turin) Fixed frequency 2.5Ghz
>> 
>> Baseline:-
>> ----------
>> Benchmark                   (shift)   Mode  Cnt      Score   Error   Units
>> ByteShiftBenchmark.ashr128        3  thrpt    2  24787.795          ops/us
>> ByteShiftBenchmark.ashr256        3  thrpt    2  29184.019          ops/us
>> ByteShiftBenchmark.ashr512        3  thrpt    2  42397.653          ops/us
>> ByteShiftBenchmark.lshl128        3  thrpt    2  25241.434          ops/us
>> ByteShiftBenchmark.lshl256        3  thrpt    2  28274.012          ops/us
>> ByteShiftBenchmark.lshl512        3  thrpt    2  47331.675          ops/us
>> ByteShiftBenchmark.lshr128        3  thrpt    2  25197.166          ops/us
>> ByteShiftBenchmark.lshr256        3  thrpt    2  27587.841          ops/us
>> ByteShiftBenchmark.lshr512        3  thrpt    2  38011.674          ops/us
>> 
>> Withopt:-
>> ---------
>> Benchmark                   (shift)   Mode  Cnt      Score   Error   Units
>> ByteShiftBenchmark.ashr128        3  thrpt    2  30674.163          ops/us
>> ByteShiftBenchmark.ashr256        3  thrpt    2  43275.297          ops/us
>> ByteShiftBenchmark.ashr512        3  thrpt    2  72744.344          ops/us
>> ByteShiftBenchmark.lshl128        3  thrpt    2  28799.612          ops/us
>> ByteShiftBenchmark.lshl256        3  thrpt    2  40797.362          ops/us
>> ByteShiftBenchmark.lshl512        3  thrpt    2  82156.329          ops/us
>> ByteShiftBenchmark.lshr128        3  thrpt    2  30325.171          ops/us
>> ByteShiftBenchmark.lshr256        3  thrpt    2  42775.756          ops/us
>> ByteShiftBenchmark.lshr512        3  thrpt    2  65255.800          ops/us
>> 
>> 
>> 
>> Kindly review and share your feedback.
>> 
>> Best Regards,
>> Jatin
>> 
>> ---------
>> - [x] I confirm that I make th...
>
> Jatin Bhateja has refreshed the contents of this pull request, and previous 
> commits have been removed. The incremental views will show differences 
> compared to the previous content of the PR. The pull request contains two new 
> commits since the last revision:
> 
>  - Rename vector_shift_gfni to vector_byte_shift_gfni
>  - 8393068: Optimize uniform vector byte shift operations using x86 GFNI 
> instruction

src/hotspot/cpu/x86/c2_MacroAssembler_x86.cpp line 1372:

> 1370: }
> 1371: 
> 1372: void C2_MacroAssembler::vshiftb_gfni(int opcode, XMMRegister dst, 
> XMMRegister src, XMMRegister shift,

Please add mathematical proof to this implementation.

-------------

PR Review Comment: https://git.openjdk.org/jdk/pull/33113#discussion_r4130371273

Reply via email to