On Mon, 28 Sep 2026 18:25:14 GMT, Chad Rakoczy <[email protected]> wrote:
>> Adds support for vectorized dot product on aarch64 (`sdot` and `udot`) >> through the Vector API >> >> Dedicated dot product instructions have been shown to provide up to 10x >> throughput for lucene >> ([results](https://github.com/apache/lucene/pull/13572)) compared to >> vectorized multiply and add. This PR adds two new methods to `ByteVector` to >> leverage the aarch dot product instructions `sdot` and `udot` respectively. >> - `IntVector dot(Vector<Byte> v, Vector<Integer> acc)` >> - `IntVector dotUnsigned(Vector<Byte> v, Vector<Integer> acc)` >> >> Each int lane of the result holds the dot product of the corresponding group >> of four bytes from the two operands, added to the matching accumulator lane. >> For example: >> >> a = [a1, a2, a3, a4, ..., ..., a13, a14, a15, a16] >> b = [b1, b2, b3, b4, ..., ..., b13, b14, b15, b16] >> acc = [acc1, ..., ..., acc4] >> >> a.dot(b, acc) -> [ >> acc1 + a1 * b1 + a2 * b2 + a3 * b3 + a4 * b4, >> ..., >> ..., >> acc4 + a13 * b13 + a14 * b14 + a15 * b15 + a16 * b16 >> ] >> >> >> The equivalent instructions on x86 are `VPDPBSSD` and `VPDPBUUD` which >> perform the same operations and match the proposed new functions however >> this PR only includes aarch64. >> >> Graviton 2 >> >> Benchmark Mode Cnt Score Error Units >> VectorDotBenchmark.dotScalar thrpt 25 2412.719 ± 0.014 ops/ms >> VectorDotBenchmark.dotMulAdd thrpt 25 1497.953 ± 2.734 ops/ms >> VectorDotBenchmark.dotVector thrpt 25 23243.016 ± 141.708 ops/ms >> VectorDotBenchmark.dotUnsignedScalar thrpt 25 2402.966 ± 0.157 ops/ms >> VectorDotBenchmark.dotUnsignedMulAdd thrpt 25 1498.898 ± 1.442 ops/ms >> VectorDotBenchmark.dotUnsignedVector thrpt 25 23344.116 ± 222.150 ops/ms >> >> >> Graviton 3 >> >> Benchmark Mode Cnt Score Error Units >> VectorDotBenchmark.dotScalar thrpt 25 7946.950 ± 2.564 ops/ms >> VectorDotBenchmark.dotMulAdd thrpt 25 3257.311 ± 5.308 ops/ms >> VectorDotBenchmark.dotVector thrpt 25 41430.996 ± 688.936 ops/ms >> VectorDotBenchmark.dotUnsignedScalar thrpt 25 2536.268 ± 0.438 ops/ms >> VectorDotBenchmark.dotUnsignedMulAdd thrpt 25 3259.853 ± 3.511 ops/ms >> VectorDotBenchmark.dotUnsignedVector thrpt 25 41263.821 ± 478.856 ops/ms >> >> >> --------- >> - [x] I confirm that I make this contribution in accordance with the >> [OpenJDK Interim AI Policy](https://openjdk.org/legal... > > @goankur @shubhamvishu This PR adds support for dot product backed by sdot > and udot on aarch64 via the Vector API. It should cover what > apache/lucene#13572 and apache/lucene#15508 are aiming to do. I'm interested > to see if you see the same performance benefits with this implementation. Any > feedback is greatly appreciated. @chadrako Looks like interesting work, especially considering the performance. However, it may go a bit against the philosophy of the Vector API, that we don't want to add all sorts of "snowflake" operations but keep the API as simple as possible. Alternative approach below, might it work? Are the semantics the same as first casting the two ByteVector to IntVector, then do the mul, then the add with the acc? If so, your computation could be expressed in current Vector API, but then just recognized by C2 during IGVN, or even better in the AD files (more platform specific). ------------- PR Comment: https://git.openjdk.org/jdk/pull/32359#issuecomment-5893487093
