On Mon, 28 Sep 2026 18:25:14 GMT, Chad Rakoczy <[email protected]> wrote:

>> Adds support for vectorized dot product on aarch64 (`sdot` and `udot`) 
>> through the Vector API
>> 
>> Dedicated dot product instructions have been shown to provide up to 10x 
>> throughput for lucene 
>> ([results](https://github.com/apache/lucene/pull/13572)) compared to 
>> vectorized multiply and add. This PR adds two new methods to `ByteVector` to 
>> leverage the aarch dot product instructions `sdot` and `udot` respectively.
>> - `IntVector dot(Vector<Byte> v, Vector<Integer> acc)`
>> - `IntVector dotUnsigned(Vector<Byte> v, Vector<Integer> acc)`
>> 
>> Each int lane of the result holds the dot product of the corresponding group 
>> of four bytes from the two operands, added to the matching accumulator lane. 
>> For example:
>> 
>> a = [a1, a2, a3, a4, ..., ..., a13, a14, a15, a16]
>> b = [b1, b2, b3, b4, ..., ..., b13, b14, b15, b16]
>> acc = [acc1, ..., ..., acc4]
>> 
>> a.dot(b, acc) -> [
>>     acc1 + a1 * b1 + a2 * b2 + a3 * b3 + a4 * b4, 
>>     ..., 
>>     ...,
>>     acc4 + a13 * b13 + a14 * b14 + a15 * b15 + a16 * b16
>> ]
>> 
>> 
>> The equivalent instructions on x86 are `VPDPBSSD` and `VPDPBUUD` which 
>> perform the same operations and match the proposed new functions however 
>> this PR only includes aarch64.
>> 
>> Graviton 2
>> 
>> Benchmark                              Mode  Cnt      Score     Error   Units
>> VectorDotBenchmark.dotScalar          thrpt   25   2412.719 ±   0.014  ops/ms
>> VectorDotBenchmark.dotMulAdd          thrpt   25   1497.953 ±   2.734  ops/ms
>> VectorDotBenchmark.dotVector          thrpt   25  23243.016 ± 141.708  ops/ms
>> VectorDotBenchmark.dotUnsignedScalar  thrpt   25   2402.966 ±   0.157  ops/ms
>> VectorDotBenchmark.dotUnsignedMulAdd  thrpt   25   1498.898 ±   1.442  ops/ms
>> VectorDotBenchmark.dotUnsignedVector  thrpt   25  23344.116 ± 222.150  ops/ms
>> 
>> 
>> Graviton 3
>> 
>> Benchmark                              Mode  Cnt      Score     Error   Units
>> VectorDotBenchmark.dotScalar          thrpt   25   7946.950 ±   2.564  ops/ms
>> VectorDotBenchmark.dotMulAdd          thrpt   25   3257.311 ±   5.308  ops/ms
>> VectorDotBenchmark.dotVector          thrpt   25  41430.996 ± 688.936  ops/ms
>> VectorDotBenchmark.dotUnsignedScalar  thrpt   25   2536.268 ±   0.438  ops/ms
>> VectorDotBenchmark.dotUnsignedMulAdd  thrpt   25   3259.853 ±   3.511  ops/ms
>> VectorDotBenchmark.dotUnsignedVector  thrpt   25  41263.821 ± 478.856  ops/ms
>> 
>> 
>> ---------
>> - [x] I confirm that I make this contribution in accordance with the 
>> [OpenJDK Interim AI Policy](https://openjdk.org/legal...
>
> @goankur @shubhamvishu This PR adds support for dot product backed by sdot 
> and udot on aarch64 via the Vector API. It should cover what 
> apache/lucene#13572 and apache/lucene#15508 are aiming to do. I'm interested 
> to see if you see the same performance benefits with this implementation. Any 
> feedback is greatly appreciated.

@chadrako Looks like interesting work, especially considering the performance. 
However, it may go a bit against the philosophy of the Vector API, that we 
don't want to add all sorts of "snowflake" operations but keep the API as 
simple as possible.

Alternative approach below, might it work?
Are the semantics the same as first casting the two ByteVector to IntVector, 
then do the mul, then the add with the acc? If so, your computation could be 
expressed in current Vector API, but then just recognized by C2 during IGVN, or 
even better in the AD files (more platform specific).

-------------

PR Comment: https://git.openjdk.org/jdk/pull/32359#issuecomment-5893487093

Reply via email to