> Hi All, > > x86 AVX-512 targets provides packed shifts for 16, 32, and 64-bit lanes, but > not for bytes. C2 lowers LShiftVB, RShiftVB, and URShiftVB by widening bytes > to words, shifting, masking, and packing. That sequence is the byte-shift > bottleneck. > > GFNI **gf2p8affineqb** instruction applies one 8×8 bit matrix to every byte. > Logical left, logical right, and arithmetic right shifts by 0..7 are fixed > matrices derived from the identity 0x0102040810204080. > > Patch uses this lowering for 128-, 256-, and 512-bit vectors when GFNI is > available and keeps the existing widen/shift/pack path when GFNI is absent. > > Following are the performance numbers of benchmark included with the patch. > > System: AMD EPYC 9755 128-Core Processor (Turin) Fixed frequency 2.5Ghz > > Baseline:- > ---------- > Benchmark (shift) Mode Cnt Score Error Units > ByteShiftBenchmark.ashr128 3 thrpt 2 24787.795 ops/us > ByteShiftBenchmark.ashr256 3 thrpt 2 29184.019 ops/us > ByteShiftBenchmark.ashr512 3 thrpt 2 42397.653 ops/us > ByteShiftBenchmark.lshl128 3 thrpt 2 25241.434 ops/us > ByteShiftBenchmark.lshl256 3 thrpt 2 28274.012 ops/us > ByteShiftBenchmark.lshl512 3 thrpt 2 47331.675 ops/us > ByteShiftBenchmark.lshr128 3 thrpt 2 25197.166 ops/us > ByteShiftBenchmark.lshr256 3 thrpt 2 27587.841 ops/us > ByteShiftBenchmark.lshr512 3 thrpt 2 38011.674 ops/us > > Withopt:- > --------- > Benchmark (shift) Mode Cnt Score Error Units > ByteShiftBenchmark.ashr128 3 thrpt 2 30674.163 ops/us > ByteShiftBenchmark.ashr256 3 thrpt 2 43275.297 ops/us > ByteShiftBenchmark.ashr512 3 thrpt 2 72744.344 ops/us > ByteShiftBenchmark.lshl128 3 thrpt 2 28799.612 ops/us > ByteShiftBenchmark.lshl256 3 thrpt 2 40797.362 ops/us > ByteShiftBenchmark.lshl512 3 thrpt 2 82156.329 ops/us > ByteShiftBenchmark.lshr128 3 thrpt 2 30325.171 ops/us > ByteShiftBenchmark.lshr256 3 thrpt 2 42775.756 ops/us > ByteShiftBenchmark.lshr512 3 thrpt 2 65255.800 ops/us > > > > Kindly review and share your feedback. > > Best Regards, > Jatin > > --------- > - [x] I confirm that I make this contribution in accordance with the [OpenJDK > Interim AI Policy](https://openjdk.org/legal/a...
Jatin Bhateja has refreshed the contents of this pull request, and previous commits have been removed. The incremental views will show differences compared to the previous content of the PR. The pull request contains two new commits since the last revision: - Rename vector_shift_gfni to vector_byte_shift_gfni - 8393068: Optimize uniform vector byte shift operations using x86 GFNI instruction ------------- Changes: - all: https://git.openjdk.org/jdk/pull/33113/files - new: https://git.openjdk.org/jdk/pull/33113/files/7dff81dc..b4d4996a Webrevs: - full: https://webrevs.openjdk.org/?repo=jdk&pr=33113&range=01 - incr: https://webrevs.openjdk.org/?repo=jdk&pr=33113&range=00-01 Stats: 9 lines in 3 files changed: 0 ins; 0 del; 9 mod Patch: https://git.openjdk.org/jdk/pull/33113.diff Fetch: git fetch https://git.openjdk.org/jdk.git pull/33113/head:pull/33113 PR: https://git.openjdk.org/jdk/pull/33113
