> Hi All,
> 
> x86 AVX-512 targets provides packed shifts for 16, 32, and 64-bit lanes, but 
> not for bytes. C2 lowers LShiftVB, RShiftVB, and URShiftVB by widening bytes 
> to words, shifting, masking, and packing. That sequence is the byte-shift 
> bottleneck.
> 
> GFNI **gf2p8affineqb** instruction applies one 8×8 bit matrix to every byte. 
> Logical left, logical right, and arithmetic right shifts by 0..7 are fixed 
> matrices derived from the identity 0x0102040810204080.
> 
> Patch uses this lowering for 128-, 256-, and 512-bit vectors when GFNI is 
> available and keeps the existing widen/shift/pack path when GFNI is absent.
> 
> Following are the performance numbers of benchmark included with the patch.
> 
> System: AMD EPYC 9755 128-Core Processor (Turin) Fixed frequency 2.5Ghz
> 
> Baseline:-
> ----------
> Benchmark                   (shift)   Mode  Cnt      Score   Error   Units
> ByteShiftBenchmark.ashr128        3  thrpt    2  24787.795          ops/us
> ByteShiftBenchmark.ashr256        3  thrpt    2  29184.019          ops/us
> ByteShiftBenchmark.ashr512        3  thrpt    2  42397.653          ops/us
> ByteShiftBenchmark.lshl128        3  thrpt    2  25241.434          ops/us
> ByteShiftBenchmark.lshl256        3  thrpt    2  28274.012          ops/us
> ByteShiftBenchmark.lshl512        3  thrpt    2  47331.675          ops/us
> ByteShiftBenchmark.lshr128        3  thrpt    2  25197.166          ops/us
> ByteShiftBenchmark.lshr256        3  thrpt    2  27587.841          ops/us
> ByteShiftBenchmark.lshr512        3  thrpt    2  38011.674          ops/us
> 
> Withopt:-
> ---------
> Benchmark                   (shift)   Mode  Cnt      Score   Error   Units
> ByteShiftBenchmark.ashr128        3  thrpt    2  30674.163          ops/us
> ByteShiftBenchmark.ashr256        3  thrpt    2  43275.297          ops/us
> ByteShiftBenchmark.ashr512        3  thrpt    2  72744.344          ops/us
> ByteShiftBenchmark.lshl128        3  thrpt    2  28799.612          ops/us
> ByteShiftBenchmark.lshl256        3  thrpt    2  40797.362          ops/us
> ByteShiftBenchmark.lshl512        3  thrpt    2  82156.329          ops/us
> ByteShiftBenchmark.lshr128        3  thrpt    2  30325.171          ops/us
> ByteShiftBenchmark.lshr256        3  thrpt    2  42775.756          ops/us
> ByteShiftBenchmark.lshr512        3  thrpt    2  65255.800          ops/us
> 
> 
> 
> Kindly review and share your feedback.
> 
> Best Regards,
> Jatin
> 
> ---------
> - [x] I confirm that I make this contribution in accordance with the [OpenJDK 
> Interim AI Policy](https://openjdk.org/legal/a...

Jatin Bhateja has refreshed the contents of this pull request, and previous 
commits have been removed. The incremental views will show differences compared 
to the previous content of the PR. The pull request contains two new commits 
since the last revision:

 - Rename vector_shift_gfni to vector_byte_shift_gfni
 - 8393068: Optimize uniform vector byte shift operations using x86 GFNI 
instruction

-------------

Changes:
  - all: https://git.openjdk.org/jdk/pull/33113/files
  - new: https://git.openjdk.org/jdk/pull/33113/files/7dff81dc..b4d4996a

Webrevs:
 - full: https://webrevs.openjdk.org/?repo=jdk&pr=33113&range=01
 - incr: https://webrevs.openjdk.org/?repo=jdk&pr=33113&range=00-01

  Stats: 9 lines in 3 files changed: 0 ins; 0 del; 9 mod
  Patch: https://git.openjdk.org/jdk/pull/33113.diff
  Fetch: git fetch https://git.openjdk.org/jdk.git pull/33113/head:pull/33113

PR: https://git.openjdk.org/jdk/pull/33113

Reply via email to