Hi All,

x86 AVX-512 targets provides packed shifts for 16, 32, and 64-bit lanes, but 
not for bytes. C2 lowers LShiftVB, RShiftVB, and URShiftVB by widening bytes to 
words, shifting, masking, and packing. That sequence is the byte-shift 
bottleneck.

GFNI **gf2p8affineqb** instruction applies one 8×8 bit matrix to every byte. 
Logical left, logical right, and arithmetic right shifts by 0..7 are fixed 
matrices derived from the identity 0x0102040810204080.

Patch uses this lowering for 128-, 256-, and 512-bit vectors when GFNI is 
available and keeps the existing widen/shift/pack path when GFNI is absent.

Following are the performance numbers of benchmark included with the patch.

System: AMD EPYC 9755 128-Core Processor (Turin) Fixed frequency 2.5Ghz

Baseline:-
----------
Benchmark                   (shift)   Mode  Cnt      Score   Error   Units
ByteShiftBenchmark.ashr128        3  thrpt    2  24787.795          ops/us
ByteShiftBenchmark.ashr256        3  thrpt    2  29184.019          ops/us
ByteShiftBenchmark.ashr512        3  thrpt    2  42397.653          ops/us
ByteShiftBenchmark.lshl128        3  thrpt    2  25241.434          ops/us
ByteShiftBenchmark.lshl256        3  thrpt    2  28274.012          ops/us
ByteShiftBenchmark.lshl512        3  thrpt    2  47331.675          ops/us
ByteShiftBenchmark.lshr128        3  thrpt    2  25197.166          ops/us
ByteShiftBenchmark.lshr256        3  thrpt    2  27587.841          ops/us
ByteShiftBenchmark.lshr512        3  thrpt    2  38011.674          ops/us

Withopt:-
---------
Benchmark                   (shift)   Mode  Cnt      Score   Error   Units
ByteShiftBenchmark.ashr128        3  thrpt    2  30674.163          ops/us
ByteShiftBenchmark.ashr256        3  thrpt    2  43275.297          ops/us
ByteShiftBenchmark.ashr512        3  thrpt    2  72744.344          ops/us
ByteShiftBenchmark.lshl128        3  thrpt    2  28799.612          ops/us
ByteShiftBenchmark.lshl256        3  thrpt    2  40797.362          ops/us
ByteShiftBenchmark.lshl512        3  thrpt    2  82156.329          ops/us
ByteShiftBenchmark.lshr128        3  thrpt    2  30325.171          ops/us
ByteShiftBenchmark.lshr256        3  thrpt    2  42775.756          ops/us
ByteShiftBenchmark.lshr512        3  thrpt    2  65255.800          ops/us



Kindly review and share your feedback.

Best Regards,
Jatin

---------
- [x] I confirm that I make this contribution in accordance with the [OpenJDK 
Interim AI Policy](https://openjdk.org/legal/ai).

-------------

Commit messages:
 - 8393068: Optimize uniform vector byte shift operations using x86 GFNI 
instruction

Changes: https://git.openjdk.org/jdk/pull/33113/files
  Webrev: https://webrevs.openjdk.org/?repo=jdk&pr=33113&range=00
  Issue: https://bugs.openjdk.org/browse/JDK-8393068
  Stats: 399 lines in 10 files changed: 395 ins; 0 del; 4 mod
  Patch: https://git.openjdk.org/jdk/pull/33113.diff
  Fetch: git fetch https://git.openjdk.org/jdk.git pull/33113/head:pull/33113

PR: https://git.openjdk.org/jdk/pull/33113

Reply via email to