On Thu, 13 Aug 2026 22:48:49 GMT, Chad Rakoczy <[email protected]> wrote:
> Adds support for vectorized dot product on aarch64 (`sdot` and `udot`) > through the Vector API > > Dedicated dot product instructions have been shown to provide up to 10x > throughput for lucene > ([results](https://github.com/apache/lucene/pull/13572)) compared to > vectorized multiply and add. This PR adds two new methods to `ByteVector` to > leverage the aarch dot product instructions `sdot` and `udot` respectively. > - `IntVector dot(Vector<Byte> v, Vector<Integer> acc)` > - `IntVector dotUnsigned(Vector<Byte> v, Vector<Integer> acc)` > > Each int lane of the result holds the dot product of the corresponding group > of four bytes from the two operands, added to the matching accumulator lane. > For example: > > a = [a1, a2, a3, a4, ..., ..., a13, a14, a15, a16] > b = [b1, b2, b3, b4, ..., ..., b13, b14, b15, b16] > acc = [acc1, ..., ..., acc4] > > a.dot(b, acc) -> [ > acc1 + a1 * b1 + a2 * b2 + a3 * b3 + a4 * b4, > ..., > ..., > acc4 + a13 * b13 + a14 * b14 + a15 * b15 + a16 * b16 > ] > > > The equivalent instructions on x86 are `VPDPBSSD` and `VPDPBUUD` which > perform the same operations and match the proposed new functions however this > PR only includes aarch64. > > Graviton 2 > > Benchmark Mode Cnt Score Error Units > VectorDotBenchmark.dotScalar thrpt 25 2412.719 ± 0.014 ops/ms > VectorDotBenchmark.dotMulAdd thrpt 25 1497.953 ± 2.734 ops/ms > VectorDotBenchmark.dotVector thrpt 25 23243.016 ± 141.708 ops/ms > VectorDotBenchmark.dotUnsignedScalar thrpt 25 2402.966 ± 0.157 ops/ms > VectorDotBenchmark.dotUnsignedMulAdd thrpt 25 1498.898 ± 1.442 ops/ms > VectorDotBenchmark.dotUnsignedVector thrpt 25 23344.116 ± 222.150 ops/ms > > > Graviton 3 > > Benchmark Mode Cnt Score Error Units > VectorDotBenchmark.dotScalar thrpt 25 7946.950 ± 2.564 ops/ms > VectorDotBenchmark.dotMulAdd thrpt 25 3257.311 ± 5.308 ops/ms > VectorDotBenchmark.dotVector thrpt 25 41430.996 ± 688.936 ops/ms > VectorDotBenchmark.dotUnsignedScalar thrpt 25 2536.268 ± 0.438 ops/ms > VectorDotBenchmark.dotUnsignedMulAdd thrpt 25 3259.853 ± 3.511 ops/ms > VectorDotBenchmark.dotUnsignedVector thrpt 25 41263.821 ± 478.856 ops/ms > > > --------- > - [x] I confirm that I make this contribution in accordance with the [OpenJDK > Interim AI Policy](https://openjdk.org/legal/ai). @goankur @shubhamvishu This PR adds support for dot product backed by sdot and udot on aarch64 via the Vector API. It should cover what apache/lucene#13572 and apache/lucene#15508 are aiming to do. I'm interested to see if you see the same performance benefits with this implementation. Any feedback is greatly appreciated. ------------- PR Comment: https://git.openjdk.org/jdk/pull/32359#issuecomment-5876019611
