> Hi All,
> 
> x86 AVX-512 targets provides packed shifts for 16, 32, and 64-bit lanes, but 
> not for bytes. C2 lowers LShiftVB, RShiftVB, and URShiftVB by widening bytes 
> to words, shifting, masking, and packing. That sequence is the byte-shift 
> bottleneck.
> 
> GFNI **gf2p8affineqb** instruction applies one 8×8 bit matrix to every byte. 
> Logical left, logical right, and arithmetic right shifts by 0..7 are fixed 
> matrices derived from the identity 0x0102040810204080.
> 
> Patch uses this lowering for 128-, 256-, and 512-bit vectors when GFNI is 
> available and keeps the existing widen/shift/pack path when GFNI is absent.
> 
> Following are the performance numbers of benchmark included with the patch.
> 
> System: AMD EPYC 9755 128-Core Processor (Turin) Fixed frequency 2.5Ghz
> 
> Baseline:-
> ----------
> Benchmark                   (shift)   Mode  Cnt      Score   Error   Units
> ByteShiftBenchmark.ashr128        3  thrpt    2  24787.795          ops/us
> ByteShiftBenchmark.ashr256        3  thrpt    2  29184.019          ops/us
> ByteShiftBenchmark.ashr512        3  thrpt    2  42397.653          ops/us
> ByteShiftBenchmark.lshl128        3  thrpt    2  25241.434          ops/us
> ByteShiftBenchmark.lshl256        3  thrpt    2  28274.012          ops/us
> ByteShiftBenchmark.lshl512        3  thrpt    2  47331.675          ops/us
> ByteShiftBenchmark.lshr128        3  thrpt    2  25197.166          ops/us
> ByteShiftBenchmark.lshr256        3  thrpt    2  27587.841          ops/us
> ByteShiftBenchmark.lshr512        3  thrpt    2  38011.674          ops/us
> 
> Withopt:-
> ---------
> Benchmark                   (shift)   Mode  Cnt      Score   Error   Units
> ByteShiftBenchmark.ashr128        3  thrpt    2  30674.163          ops/us
> ByteShiftBenchmark.ashr256        3  thrpt    2  43275.297          ops/us
> ByteShiftBenchmark.ashr512        3  thrpt    2  72744.344          ops/us
> ByteShiftBenchmark.lshl128        3  thrpt    2  28799.612          ops/us
> ByteShiftBenchmark.lshl256        3  thrpt    2  40797.362          ops/us
> ByteShiftBenchmark.lshl512        3  thrpt    2  82156.329          ops/us
> ByteShiftBenchmark.lshr128        3  thrpt    2  30325.171          ops/us
> ByteShiftBenchmark.lshr256        3  thrpt    2  42775.756          ops/us
> ByteShiftBenchmark.lshr512        3  thrpt    2  65255.800          ops/us
> 
> 
> 
> Kindly review and share your feedback.
> 
> Best Regards,
> Jatin
> 
> ---------
> - [x] I confirm that I make this contribution in accordance with the [OpenJDK 
> Interim AI Policy](https://openjdk.org/legal/a...

Jatin Bhateja has updated the pull request incrementally with one additional 
commit since the last revision:

  Review comments resolution

-------------

Changes:
  - all: https://git.openjdk.org/jdk/pull/33113/files
  - new: https://git.openjdk.org/jdk/pull/33113/files/b4d4996a..67b4424e

Webrevs:
 - full: https://webrevs.openjdk.org/?repo=jdk&pr=33113&range=02
 - incr: https://webrevs.openjdk.org/?repo=jdk&pr=33113&range=01-02

  Stats: 45 lines in 1 file changed: 45 ins; 0 del; 0 mod
  Patch: https://git.openjdk.org/jdk/pull/33113.diff
  Fetch: git fetch https://git.openjdk.org/jdk.git pull/33113/head:pull/33113

PR: https://git.openjdk.org/jdk/pull/33113

Reply via email to