On Wed, Sep 30, 2026 at 01:19:25AM +0530, Amit Machhiwal wrote:
> Hi Michal,
>
> Thanks for sharing the host information and the full guest boot log.
>
> On 2026/09/29 02:18 PM, Michal Suchánek wrote:
> > On Tue, Sep 29, 2026 at 04:39:05PM +0530, Amit Machhiwal wrote:
> > > On 2026/09/29 01:04 PM, Michal Suchánek wrote:
> > > > On Tue, Sep 29, 2026 at 09:36:05AM +0530, Amit Machhiwal wrote:
> > > > > On 2026/09/28 11:43 AM, Amit Machhiwal wrote:
> > > > > > Hi Michal,
> > > > > >
> > > > > > Thanks for the report. We have been looking at the crash and have
> > > > > > some findings
> > > > > > and follow-up questions.
> > > > > >
> > > > > > On 2026/09/25 10:10 AM, Michal Suchánek wrote:
> > > > > > > Hello,
> > > > > > >
> > > > > > > There appears to be a regression between
> > > > >
> > > > > < snip >
> > > > >
> > > > > > >
> > > > > > > Linux 7.0.12
> > > > > > > https://github.com/openSUSE/kernel/tree/4c9922875554ec191166ba6c0145f12d8ff9d4fc
> > > > > > > https://github.com/openSUSE/kernel-source/blob/2ebf0bc/config/ppc64le/default
> > > > >
> > > > > I have been trying to recreate the issue with this kernel and config
> > > > > but I
> > > > > haven't been able to. The L2 KVM guest boots fine everytime.
> > > > >
> > > > > In addition to the requested information, could you please also share
> > > > > your qemu
> > > > > cmdline/guest xml you used?
> > > >
> > > > The XML of the VM is below. What other information do you require?
> > >
> > > Thanks for sharing the guest XML. I had requested some more info [1].
> > > Could
> > > you please share that?
> > >
> > > [1]
> > > https://lore.kernel.org/all/[email protected]/
> >
> > Hello,
> >
> > full log below.
> >
> > host:
> >
> > numactl --hardware
> > available: 2 nodes (0-1)
> > node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
> > 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48
> > 49 50 51 52 53 54 55
> > node 0 size: 122139 MB
> > node 0 free: 103635 MB
> > node 1 cpus: 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76
> > 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100
> > 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119
> > node 1 size: 139516 MB
> > node 1 free: 128483 MB
> > node distances:
> > node 0 1
> > 0: 10 20
> > 1: 20 10
> >
> > The patch does indeed make it possible to boot when the buffer
> > allocation fails.
>
> Glad the first fix works. That fixes the crash (the NULL deref in
> __find_event_file()), but the ring buffer allocation failure itself is a
> separate bug worth fixing too.
>
> My suspicion is that with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y,
> si_mem_available()
> returns falsely negative during early boot because NR_FREE_PAGES only reflects
> the non-deferred memory pool at that point — the bulk of RAM hasn't been
> handed
> to the buddy allocator yet. The check in __rb_allocate_pages() rejects the
> allocation prematurely.
>
> [ 0.149312][ T743] node 0 deferred pages initialised in 20ms
>
> Since the issue is not recreating on my environment, could you please give
> this
> patch a try and see if the actual crash goes away?
>
> diff --git a/kernel/trace/ring_buffer.c b/kernel/trace/ring_buffer.c
> index 04bb94c29f58..224cc0e5c066 100644
> --- a/kernel/trace/ring_buffer.c
> +++ b/kernel/trace/ring_buffer.c
> @@ -2454,7 +2454,7 @@ static int __rb_allocate_pages(struct
> ring_buffer_per_cpu *cpu_buffer,
> * not going to succeed.
> */
> i = si_mem_available();
> - if (i < nr_pages)
> + if (system_state != SYSTEM_BOOTING && i < nr_pages)
> return -ENOMEM;
>
> /*
Hello,
this does seem to mitigate the problem with allocating the trace buffer.
Related to reproducing the problem it seems that it is much more likely
to happen on reboot than first boot, and the typical solution for kernel
test automation produces one-off VMs that never reboot. With the VM XML
as posted earlier using a full production OS image testing is not very
efficient. Still with something like
n=1 ; while true ; do echo $n ; n=$(expr $n + 1) ; ssh 192.168.0.2
reboot ; sleep 10 ; ssh 192.168.0.2 dmesg | { grep ERROR: && exit 1 ;
} ; done
the VM can be rebooted a number of times without additional user
interaction.
Thanks
Michal