On Thu, 13 Aug 2026 22:48:49 GMT, Chad Rakoczy <[email protected]> wrote:

> Adds support for vectorized dot product on aarch64 (`sdot` and `udot`) 
> through the Vector API
> 
> Dedicated dot product instructions have been shown to provide up to 10x 
> throughput for lucene 
> ([results](https://github.com/apache/lucene/pull/13572)) compared to 
> vectorized multiply and add. This PR adds two new methods to `ByteVector` to 
> leverage the aarch dot product instructions `sdot` and `udot` respectively.
> - `IntVector dot(Vector<Byte> v, Vector<Integer> acc)`
> - `IntVector dotUnsigned(Vector<Byte> v, Vector<Integer> acc)`
> 
> Each int lane of the result holds the dot product of the corresponding group 
> of four bytes from the two operands, added to the matching accumulator lane. 
> For example:
> 
> a = [a1, a2, a3, a4, ..., ..., a13, a14, a15, a16]
> b = [b1, b2, b3, b4, ..., ..., b13, b14, b15, b16]
> acc = [acc1, ..., ..., acc4]
> 
> a.dot(b, acc) -> [
>     acc1 + a1 * b1 + a2 * b2 + a3 * b3 + a4 * b4, 
>     ..., 
>     ...,
>     acc4 + a13 * b13 + a14 * b14 + a15 * b15 + a16 * b16
> ]
> 
> 
> The equivalent instructions on x86 are `VPDPBSSD` and `VPDPBUUD` which 
> perform the same operations and match the proposed new functions however this 
> PR only includes aarch64.
> 
> Graviton 2
> 
> Benchmark                              Mode  Cnt      Score     Error   Units
> VectorDotBenchmark.dotScalar          thrpt   25   2412.719 ±   0.014  ops/ms
> VectorDotBenchmark.dotMulAdd          thrpt   25   1497.953 ±   2.734  ops/ms
> VectorDotBenchmark.dotVector          thrpt   25  23243.016 ± 141.708  ops/ms
> VectorDotBenchmark.dotUnsignedScalar  thrpt   25   2402.966 ±   0.157  ops/ms
> VectorDotBenchmark.dotUnsignedMulAdd  thrpt   25   1498.898 ±   1.442  ops/ms
> VectorDotBenchmark.dotUnsignedVector  thrpt   25  23344.116 ± 222.150  ops/ms
> 
> 
> Graviton 3
> 
> Benchmark                              Mode  Cnt      Score     Error   Units
> VectorDotBenchmark.dotScalar          thrpt   25   7946.950 ±   2.564  ops/ms
> VectorDotBenchmark.dotMulAdd          thrpt   25   3257.311 ±   5.308  ops/ms
> VectorDotBenchmark.dotVector          thrpt   25  41430.996 ± 688.936  ops/ms
> VectorDotBenchmark.dotUnsignedScalar  thrpt   25   2536.268 ±   0.438  ops/ms
> VectorDotBenchmark.dotUnsignedMulAdd  thrpt   25   3259.853 ±   3.511  ops/ms
> VectorDotBenchmark.dotUnsignedVector  thrpt   25  41263.821 ± 478.856  ops/ms
> 
> 
> ---------
> - [x] I confirm that I make this contribution in accordance with the [OpenJDK 
> Interim AI Policy](https://openjdk.org/legal/ai).

@goankur @shubhamvishu This PR adds support for dot product backed by sdot and 
udot on aarch64 via the Vector API. It should cover what apache/lucene#13572 
and apache/lucene#15508 are aiming to do. I'm interested to see if you see the 
same performance benefits with this implementation. Any feedback is greatly 
appreciated.

-------------

PR Comment: https://git.openjdk.org/jdk/pull/32359#issuecomment-5876019611

Reply via email to