On Tue, 5 May 2026 12:46:56 GMT, Paul Hübner <[email protected]> wrote:

> Hi all,
> 
> This patch fixes some performance issues with AArch64's `NMethodEntryBarrier`.
> 
> The purpose of a method entrypoint is to act as a barrier for potential GC 
> work. Entrypoints can be "armed", such that when the method is called, a 
> slow-path is triggered. This slow-path then does any necessary fixups. 
> 
> With the introduction of scalarized calling conventions in Valhalla, methods 
> have multiple entrypoints, depending on the calling conventions that are 
> used. This is due to the fact that `nmethod`s cannot be multiplexed when 
> scalarized. There can be up to three entrypoints for a given method. The 
> locations of these entrypoints is tracked via relocations. 
> 
> The problem lies within the usage of `RelocIterator`. When `static void 
> set_value` is called (such as when arming entrypoints), it creates a 
> `NativeNMethodBarrier` for each entrypoint. However, the `RelocIterator` only 
> finds the first relocation, meaning each embodiment of `NativeNMethodBarrier` 
> sees the same `_guard_addr`. Consequently, executions entering through the 
> other entrypoints will not see the up-to-date armed value, and be 
> unnecessarily slow-pathed. 
> 
> This patch introduces several improvements:
> 1. Makes `NativeNMethodBarrier` use the `ldr` instructions to determine the 
> `_guard_addr`s, as opposed to relocations (contributed by @TobiHartmann). 
> This gets the correct addresses for each of the guards, and avoids having to 
> iterate through the `nmethod`'s relocations, which is expensive for large 
> methods. 
> 2. Consolidates all the entrypoints as first-class citizens of 
> `NativeNMethodBarrier`, instead of having an `NativeNMethodBarrier` per 
> entrypoint. This reduces object allocation, and makes the code much easier to 
> read for free! 
> 3. Stricter assertions. All entrypoints `ldr` instructions are verified (as 
> opposed to just the scalarized ones previously). All guard addresses are 
> verified to be within the `nmethod`'s instructions/stubs.
> 
> Testing: tiers 1-4 and some adhoc testing with `--enable-preview`. 
> 
> ---------
> - [x] I confirm that I make this contribution in accordance with the [OpenJDK 
> Interim AI Policy](https://openjdk.org/legal/ai).

This looks good to me as a point fix for the performance regression. I agree 
with @dean-long and @jsikstro though that it would be nice to optimize this to 
only use one guard value, but I think this can be done as a follow-up 
enhancement.

src/hotspot/cpu/aarch64/gc/shared/barrierSetNMethod_aarch64.cpp line 84:

> 82: // (neither instruction nor guard are nullptr).
> 83: //
> 84: // When using C1/C2 and scalarization, up to two additional (verified) 
> entrypoints,

Suggestion:

// When using the scalarized calling convention, up to two additional 
(verified) entrypoints,

-------------

Marked as reviewed by thartmann (Committer).

PR Review: 
https://git.openjdk.org/valhalla/pull/2389#pullrequestreview-4242122839
PR Review Comment: 
https://git.openjdk.org/valhalla/pull/2389#discussion_r3199801076

Reply via email to