On Fri, Sep 18, 2026 at 10:26 AM Guilherme G. Piccoli
<[email protected]> wrote:
>
> Hi Petr, Zack - thanks for CCing me!
> Some comments below:

Hi, Guilherme.

Thanks for looking at this and for adding Stephen. I'll keep you both copied.

> On 18/09/2026 00:23, Zack Rusin wrote:
> >> [...]
> >> Maybe, we should start with something simple, and introduce
> >> one more panic notifier as a start. It might be called either:
> >>
> >>   + "panic_hypervisor_list" because "crash_kexec_post_notifiers = true"
> >>     seems to be primary set on hypervisors.
> >>
> >> But I would rather make it more generic and call it
> >>
> >>   + panic_pre_crash_kexec or panic_pre_kdump because there might be
> >>     more notifiers which are either 100% safe and useful or are worth
> >>     the risk before calling crash dump.
> >>
> >> We could put there x86/vmware notifiers as a start. And we could later
> >> move there other important notifiers.
> >>
> >> How does that sound, please?
> >
>
> It's a good idea, IMO. We could start with this, Zach commented some
> implementation details below...and after it gets merged, we could move
> other hypervisors that currently set "crash_kexec_post_notifiers" to
> this list and eventually, unexport this symbol. We should avoid having
> code forcing this parameter, as Petr said, many notifiers are executed
> if that is set.

Agreed. The v2 I'm working on drops VMware's assignment to
crash_kexec_post_notifiers and leaves the existing setting unchanged.

> (I'm CCing Stephen Brennan here, I recall he had problems with this
> being auto-set, we talked about that in the panic notifiers big
> discussions in the past heh)
>
> The only thing I'd like to suggest: I think we should have a parameter
> that disables running this list, which would be the opposite of
> "crash_kexec_post_notifiers".
>
> I would implement it as something like: "postpone_pre_kexec_notifiers"
> or something like that. The parameter would basically "move" this list
> execution to the same time as the current notifiers, gating them to
> "crash_kexec_post_notifiers". This way, we'd allow users to debug kexec
> failures maybe related to the "early" notifiers. WDYT?

I'm happy to add that as a separate patch if Petr agrees. With it set,
the new list would follow ordinary panic-notifier ordering relative to
kdump: it would run before a successful transition only when
crash_kexec_post_notifiers is also set. If panic reaches the late
site, the list would remain eligible to run there. I'd keep that site
after sys_info() and before the kmsg dumpers so the log includes the
additional notifier and panic_print output available at that point.

I've called it panic_pre_kdump_postpone after the list, but I'm fine
with whatever name you and Petr prefer.

> > [...]
> > I think that without that default though, x86 oops_end() can enter
> > crash_kexec(regs) before reaching panic(), for example with
> > panic_on_oops=1. To cover that path too, I'd call the chain from
> > __crash_kexec() after the image check and register capture, under the
> > existing kexec lock. A second call in vpanic(), immediately before
> > kmsg_dump_desc(), would cover the fallback path. And I think a
> > set-once guard would prevent duplicate or recursive dispatch.
> >
>
> Regarding this, 2 things:
>
> a) I think you could change kexec_should_crash() to "return 0" also in
> case the new list is set to run, the same is done currently for
> "crash_kexec_post_notifiers". Makes sense?

I'd prefer to leave kexec_should_crash() unchanged. The direct oops
path supplies the exception registers to crash_kexec(regs), while
routing it through panic() would capture later state instead. In my
early v2 tests, the vmcores from the direct-oops path retain the
original fault registers in the crash notes. Some crash callers, for
example uv_nmi_kdump(), also bypass kexec_should_crash(). Calling the
chain from __crash_kexec() after register capture covers those paths
without changing their routing, and the shared once-only guard
prevents duplicate or recursive dispatch.

> b) Well, does this whole panic diag thing you're implementing here aims
> only at x86 guests ? Or would it be possible to run, for example, arm64
> guests? Asking this because in x86 and some other architectures (but not
> arm64[0]), it's possible to override machine_crash_shutdown() handler,
> and run things prior to a kexec. Take a look on how Hyper-V does that on
> arch/x86 - this could be just what you need, except if you plan to have
> it for all architectures heh

I'd like to support arm64 guests in the near future, but I figured
especially for review sake to limit our client in this series to x86.
So I prefer the common chain Petr proposed: as the thread you linked
shows, arm64 deliberately has no such override, and the chain gives
other clients a place to migrate away from forcing
crash_kexec_post_notifiers.

Would keeping the two call sites, with postponement in a separate
patch, work for you and Petr?

z

Attachment: smime.p7s
Description: S/MIME Cryptographic Signature

Reply via email to