Hi All, x86 AVX-512 targets provides packed shifts for 16, 32, and 64-bit lanes, but not for bytes. C2 lowers LShiftVB, RShiftVB, and URShiftVB by widening bytes to words, shifting, masking, and packing. That sequence is the byte-shift bottleneck.
GFNI **gf2p8affineqb** instruction applies one 8×8 bit matrix to every byte. Logical left, logical right, and arithmetic right shifts by 0..7 are fixed matrices derived from the identity 0x0102040810204080. Patch uses this lowering for 128-, 256-, and 512-bit vectors when GFNI is available and keeps the existing widen/shift/pack path when GFNI is absent. Following are the performance numbers of benchmark included with the patch. System: AMD EPYC 9755 128-Core Processor (Turin) Fixed frequency 2.5Ghz Baseline:- ---------- Benchmark (shift) Mode Cnt Score Error Units ByteShiftBenchmark.ashr128 3 thrpt 2 24787.795 ops/us ByteShiftBenchmark.ashr256 3 thrpt 2 29184.019 ops/us ByteShiftBenchmark.ashr512 3 thrpt 2 42397.653 ops/us ByteShiftBenchmark.lshl128 3 thrpt 2 25241.434 ops/us ByteShiftBenchmark.lshl256 3 thrpt 2 28274.012 ops/us ByteShiftBenchmark.lshl512 3 thrpt 2 47331.675 ops/us ByteShiftBenchmark.lshr128 3 thrpt 2 25197.166 ops/us ByteShiftBenchmark.lshr256 3 thrpt 2 27587.841 ops/us ByteShiftBenchmark.lshr512 3 thrpt 2 38011.674 ops/us Withopt:- --------- Benchmark (shift) Mode Cnt Score Error Units ByteShiftBenchmark.ashr128 3 thrpt 2 30674.163 ops/us ByteShiftBenchmark.ashr256 3 thrpt 2 43275.297 ops/us ByteShiftBenchmark.ashr512 3 thrpt 2 72744.344 ops/us ByteShiftBenchmark.lshl128 3 thrpt 2 28799.612 ops/us ByteShiftBenchmark.lshl256 3 thrpt 2 40797.362 ops/us ByteShiftBenchmark.lshl512 3 thrpt 2 82156.329 ops/us ByteShiftBenchmark.lshr128 3 thrpt 2 30325.171 ops/us ByteShiftBenchmark.lshr256 3 thrpt 2 42775.756 ops/us ByteShiftBenchmark.lshr512 3 thrpt 2 65255.800 ops/us Kindly review and share your feedback. Best Regards, Jatin --------- - [x] I confirm that I make this contribution in accordance with the [OpenJDK Interim AI Policy](https://openjdk.org/legal/ai). ------------- Commit messages: - 8393068: Optimize uniform vector byte shift operations using x86 GFNI instruction Changes: https://git.openjdk.org/jdk/pull/33113/files Webrev: https://webrevs.openjdk.org/?repo=jdk&pr=33113&range=00 Issue: https://bugs.openjdk.org/browse/JDK-8393068 Stats: 399 lines in 10 files changed: 395 ins; 0 del; 4 mod Patch: https://git.openjdk.org/jdk/pull/33113.diff Fetch: git fetch https://git.openjdk.org/jdk.git pull/33113/head:pull/33113 PR: https://git.openjdk.org/jdk/pull/33113
