On Tue, 29 Sep 2026 15:37:17 GMT, Emanuel Peter <[email protected]> wrote:
>> @goankur @shubhamvishu This PR adds support for dot product backed by sdot >> and udot on aarch64 via the Vector API. It should cover what >> apache/lucene#13572 and apache/lucene#15508 are aiming to do. I'm interested >> to see if you see the same performance benefits with this implementation. >> Any feedback is greatly appreciated. > > @chadrako Looks like interesting work, especially considering the > performance. However, it may go a bit against the philosophy of the Vector > API, that we don't want to add all sorts of "snowflake" operations but keep > the API as simple as possible. > > Alternative approach below, might it work? > Are the semantics the same as first casting the two ByteVector to IntVector, > then do the mul, then the add with the acc? If so, your computation could be > expressed in current Vector API, but then just recognized by C2 during IGVN, > or even better in the AD files (more platform specific). You could match the > C2 vector ops, and directly emit the required dot-product assembly > instruction. > > Ah no, it is more complicated than just mul and add. There is indeed a > reduction here. Yeah, looks like a snowflake operation to me. Matching this > in C2 IGVN/AD would be tricky, maybe fragile too. @eme64 > Ah no, it is more complicated than just mul and add. There is indeed a > reduction here. Yeah, looks like a snowflake operation to me. Matching this > in C2 IGVN/AD would be tricky, maybe fragile too. This was the motivation for the new API. The instruction `sdot` takes in 16 bytes per operand and accumulates into 4 ints. The built in group reduction makes this difficult to express with the current API. > However, it may go a bit against the philosophy of the Vector API, that we > don't want to add all sorts of "snowflake" operations but keep the API as > simple as possible. Do you think dot product as an operation does not fit into the philosophy of the Vector API? Or is it more about the current API structure? Assuming the latter, what do you think about making the API hardware agnostic by returning a scalar? Instead the API would be `int dot(Vector<Byte> v)`. This would remove the `sdot` specific API structure such as the `IntVector` result and provided accumulator. We would need to reduce every iteration, but this would allow us to use `sdot` which I still think would be an improvement. I suspect it's possible to move the reduction out of the loop so we could see near identical performance gains. ------------- PR Comment: https://git.openjdk.org/jdk/pull/32359#issuecomment-6042595061
