On Wed, 19 Aug 2026 05:52:40 GMT, Shawn Emery <[email protected]> wrote:

>> This enhancement provides AArch64 GPR intrinsics for doubleKeccak().  
>> Previously, only SIMD (Neon) intrinsics were implemented for doubleKeccak() 
>> on AArch64 systems.  Performance gains for ML-KEM and ML-DSA benchmarks 
>> improve from 2 to 9% with the GPR intrinsics:
>> 
>> ML-KEM decapsulation: +2-6% ops/sec
>> ML-KEM encapsulation: +3-8% ops/sec
>> ML-KEM key generation: +4-6% ops/sec
>> 
>> ML-DSA signing: +2-4% ops/sec
>> ML-DSA verification: +6-9% ops/sec
>> ML-DSA key generation: +6-8% ops/sec
>> 
>> ---------
>> - [X] I confirm that I make this contribution in accordance with the 
>> [OpenJDK Interim AI Policy](https://openjdk.org/legal/ai).
>
> Shawn Emery has updated the pull request incrementally with one additional 
> commit since the last revision:
> 
>   Implement comments from theRealAph and adinn

src/hotspot/cpu/aarch64/stubGenerator_aarch64.cpp line 5070:

> 5068:     }
> 5069:     __ str(a[24], Address(state, 192));
> 5070:   }

Suggestion:

    int i;
    for (i = 0; i < 24; i += 2) {
      __ stp(a[i], a[i + 1], Address(state, i * wordSize));
    }
    __ str(a[i], Address(state, i * wordSize));

-------------

PR Review Comment: https://git.openjdk.org/jdk/pull/32049#discussion_r3815335328

Reply via email to