Hi James, On Wed, Sep 30, 2026 at 04:28:37PM +0100, James Clark wrote: > If you enable pseudo-NMIs, watchdog_hardlockup_enable() installs the > watchdog using a pinned PMU event. If host userspace also has a pinned event > on the mandatory 1 PMU counter assigned to the host, then a guest could > potentially squeeze out the watchdog. > > I'm wondering if we need to prioritise kernel owned events? Or we just treat > them the same as any other event, and with PMU partitioning assume they > can't be guaranteed to be running? I feel like you would expect a watchdog > to be a bit more than best effort though, especially if there was always a > guaranteed counter available to put it on. > > I didn't follow it through completely, but it also looks like if the event > gets squeezed it would enter an error state and then never be re-enabled, > even after the guest stops running. > > Note, that I think the current ordering means that the watchdog won't > actually get squeezed out because it's created first. But I don't think we > can rely on the ordering as a strong guarantee, and it might get broken by > refactoring in the future.
Good catch. Even without future refactoring, creation order can be defeated today if /proc/sys/kernel/nmi_watchdog is toggled off and back on at runtime while a userspace per-CPU pinned event is already open, giving the re-created watchdog a higher group_index and causing merge_sched_in() to put it into PERF_EVENT_STATE_ERROR when a partitioned guest takes PMCCNTR_EL0. Prioritizing kernel-owned events (!event->owner) ahead of userspace events in perf_event_groups_cmp() and perf_less_group_idx() guarantees that a kernel-pinned watchdog always wins the mandatory host counter over userspace pinned events regardless of creation order. I will look into adding that for v10. Thanks, Colton

