On Fri, Sep 11, 2026 at 10:39:10PM +0530, Shrikanth Hegde wrote:
> Hi Michal,
> 
> On 9/10/26 5:46 PM, Michal Suchánek wrote:
> > On Thu, Sep 10, 2026 at 01:29:08PM +0200, Michal Suchánek wrote:
> > > On Thu, Sep 10, 2026 at 04:11:20PM +0530, Shrikanth Hegde wrote:
> > > > 
> > > > 
> > > > On 9/10/26 2:29 PM, Michal Suchánek wrote:
> > > > > On Fri, Sep 04, 2026 at 02:34:21PM +0530, Shrikanth Hegde wrote:
> > > > > > 
> > > > > > 
> > > > > > On 9/4/26 1:05 PM, Michal Suchánek wrote:
> > > > > > > On Thu, Sep 03, 2026 at 10:52:59PM +0530, Shrikanth Hegde wrote:
> > > > > > > > 
> > > > > > > > > > If possible run against current upstream and share the 
> > > > > > > > > > results.
> > > > > > > > > 
> > > > > > > > > https://github.com/openSUSE/kernel-source/blob/6824496d1801f73def615dca8794202eeb7b0d86/config/ppc64le/default
> > > > > > > > > 
> > > > > > > > > [  472.091531][ T6181] Kernel panic - not syncing: 
> > > > > > > > > stack-protector: Kernel stack is corrupted in: 
> > > > > > > > > kvmhv_run_single_vcpu+0x19d4/0x1b50 [kvm_hv]
> > > > > > > > > [  472.091598][ T6181] CPU: 29 UID: 107 PID: 6181 Comm: CPU 
> > > > > > > > > 112/KVM Not tainted 7.2.2-5.g6824496-default #1 PREEMPT(full) 
> > > > > > > > > openSUSE Tumbleweed (unreleased)  
> > > > > > > > > 61871a5f06863b4006f5ef27cd9f18e8a7a2edad
> > > > > > > > > [  472.091612][ T6181] Hardware name: IBM,9824-42A Power11 
> > > > > > > > > (architected) 0x820200 0xf000007 of:IBM,FW1110.20 
> > > > > > > > > (OB1110_130) hv:phyp pSeries
> > > > > > > > > [  472.091624][ T6181] Call Trace:
> > > > > > > > > [  472.091630][ T6181] [c00000000fdfb680] [c00000000134ce90] 
> > > > > > > > > dump_stack_lvl+0x84/0xc0 (unreliable)
> > > > > > > > > [  472.091653][ T6181] [c00000000fdfb6b0] [c00000000022e7c8] 
> > > > > > > > > vpanic+0x324/0x5e4
> > > > > > > > > [  472.091666][ T6181] [c00000000fdfb760] [c00000000022eac4] 
> > > > > > > > > do_panic_on_target_cpu+0x0/0x2c
> > > > > > > > > [  472.091677][ T6181] [c00000000fdfb780] [c0000000013c1ff8] 
> > > > > > > > > __stack_chk_fail+0x48/0x60
> > > > > > > > > [  472.091689][ T6181] [c00000000fdfb7f0] [c00800001aae219c] 
> > > > > > > > > kvmhv_run_single_vcpu+0x19d4/0x1b50 [kvm_hv]
> > > > > > > > > [  472.091712][ T6181] [c00000000fdfb940] [c00800001aae24b4] 
> > > > > > > > > kvmppc_vcpu_run_hv+0x19c/0x12f0 [kvm_hv]
> > > > > > > > > [  472.091732][ T6181] [c00000000fdfba10] [c00800001aeeed18] 
> > > > > > > > > kvmppc_vcpu_run+0x30/0x48 [kvm]
> > > > > > > > > [  472.091779][ T6181] [c00000000fdfba30] [c00800001aee9ef4] 
> > > > > > > > > kvm_arch_vcpu_ioctl_run+0x35c/0x4a0 [kvm]
> > > > > > > > > [  472.091813][ T6181] [c00000000fdfbac0] [c00800001aedaac4] 
> > > > > > > > > kvm_vcpu_ioctl+0x1ac/0xad8 [kvm]
> > > > > > > > > [  472.091844][ T6181] [c00000000fdfbca0] [c0000000007f1244] 
> > > > > > > > > sys_ioctl+0x374/0x1060
> > > > > > > > > [  472.091857][ T6181] [c00000000fdfbdb0] [c00000000002f7f8] 
> > > > > > > > > system_call_exception+0x188/0x430
> > > > > > > > > [  472.091871][ T6181] [c00000000fdfbe50] [c00000000000cfdc] 
> > > > > > > > > system_call_vectored_common+0x15c/0x2ec
> > > > > > > > > [  472.091886][ T6181] ---- interrupt: 3000 at 0x7fffb5565fac
> > > > > > > > > [  472.091896][ T6181] NIP:  00007fffb5565fac LR: 
> > > > > > > > > 00007fffb5565fac CTR: 0000000000000000
> > > > > > > > > [  472.091904][ T6181] REGS: c00000000fdfbe80 TRAP: 3000   
> > > > > > > > > Not tainted  (7.2.2-5.g6824496-default)
> > > > > > > > > [  472.091911][ T6181] MSR:  800000000280f033 
> > > > > > > > > <SF,VEC,VSX,EE,PR,FP,ME,IR,DR,RI,LE>  CR: 42044402  XER: 
> > > > > > > > > 00000000
> > > > > > > > > [  472.091938][ T6181] IRQMASK: 0
> > > > > > > > > [  472.091938][ T6181] GPR00: 0000000000000036 
> > > > > > > > > 00007fbfa77ed7a0 00007fffb5677100 00000000000000fa
> > > > > > > > > [  472.091938][ T6181] GPR04: 000000002000ae80 
> > > > > > > > > 0000000000000000 0000000000000000 0000000000000000
> > > > > > > > > [  472.091938][ T6181] GPR08: 00000000000000fa 
> > > > > > > > > 0000000000000000 0000000000000000 0000000000000000
> > > > > > > > > [  472.091938][ T6181] GPR12: 0000000000000000 
> > > > > > > > > 00007fbfa77f5ec0 000000014676f000 00007fbfa77ee7c0
> > > > > > > > > [  472.091938][ T6181] GPR16: 000000014674e8d0 
> > > > > > > > > 00007fbfa77eeec0 00007fbfa77eeec0 fffffffffffffff7
> > > > > > > > > [  472.091938][ T6181] GPR20: 00007fffb71210d0 
> > > > > > > > > 0000000000000001 00007fbfa77eeec0 0000000000000000
> > > > > > > > > [  472.091938][ T6181] GPR24: 00007fbfa77ed8e8 
> > > > > > > > > 0000000105971428 000000002000ae80 0000000105f77a70
> > > > > > > > > [  472.091938][ T6181] GPR28: 0000000000000000 
> > > > > > > > > 0000000000000000 000000002000ae80 000000014674f000
> > > > > > > > > [  472.092020][ T6181] NIP [00007fffb5565fac] 0x7fffb5565fac
> > > > > > > > > [  472.092027][ T6181] LR [00007fffb5565fac] 0x7fffb5565fac
> > > > > > > > > [  472.092033][ T6181] ---- interrupt: 3000
> > > > > > > > > [  472.098256][ T6181] pstore: backend (nvram) writing error 
> > > > > > > > > (-1)
> > > > > > > > > 
> > > > > > > > > This is the host, cannot run the kernel as guest because it 
> > > > > > > > > fails to boot most
> > > > > > > > > of the time inside KVM.
> > > > > > > > > 
> > > > > > > > > > 
> > > > > > > > > > > Nonethless, there are quite a few platforms. Originally 
> > > > > > > > > > > no preemption
> > > > > > > > > > > was the only option, and that's the reason why many 
> > > > > > > > > > > people run that.
> > > > > > > > > > > It's the conservative, known working option. And that's 
> > > > > > > > > > > the reason a lot
> > > > > > > > > > > of platfrom code does not get tested with more aggressive 
> > > > > > > > > > > preemtion
> > > > > > > > > > > models, and never gets fixed to work with them.
> > > > > > > > > > > 
> > > > > > > > > > 
> > > > > > > > > > Full preemption has been there for many years!.
> > > > > > > > > 
> > > > > > > > > Possible for years, forced only recently.
> > > > > > > > > 
> > > > > > > > > > Lazy is not that aggressive compared to that.
> > > > > > > > > > 
> > > > > > > > > > > Simply disabling the no preemtion option does not make 
> > > > > > > > > > > the platform code
> > > > > > > > > > > ready.
> > > > > > > > > > > 
> > > > > > > > > > 
> > > > > > > > > > Let's understand your crash case. Let's see where it is 
> > > > > > > > > > going wrong. I am suspecting
> > > > > > > > > > it is some wrong usage of preemption api rather than arch 
> > > > > > > > > > can't support preemption.
> > > > > > > > > 
> > > > > > > > > Very likely some wrong use of the preemption API by the arch 
> > > > > > > > > code, or no
> > > > > > > > > use where it should have been used. It did not matter so long 
> > > > > > > > > as people
> > > > > > > > > could run their no preempt configs and ignore the problem.
> > > > > > > > > 
> > > > > > > > > Thanks
> > > > > > > > 
> > > > > > > > 
> > > > > > > > Friendly LLM analysis says preemption is enabled too early and 
> > > > > > > > before
> > > > > > > > completing kvmppc_handle_exit_hv() / kvmppc_handle_nested_exit()
> > > > > > > > 
> > > > > > > > Below is ONLY a speculation and completely UNTESTED.
> > > > > > > > Maybe worth a try.
> > > > > > > 
> > > > > > > The patch is munged by the e-mail client, and it causes immediate
> > > > > > > voluntary preemprion in rcu critical section and hard lockup on 
> > > > > > > starting
> > > > > > > a KVM VM.
> > > > > > > 
> > > > > > > Also it would be sort of bad news if it worked because that would 
> > > > > > > be
> > > > > > > specific to book3s KVM HV and would not help with the KVM HV from 
> > > > > > > the
> > > > > > > original report which likely is not book3s, nor with KVM PR.
> > > > > > > 
> > > > > > 
> > > > > > Thanks for trying. We will try a local repro and look into it why 
> > > > > > stack is
> > > > > > getting corrupted.
> > > > > 
> 
> As we discussed offlist, samir helped to run a similar test on his machine, 
> and he didn't
> run into issue so far. we will try more.
> 
> I am just wondering what different in your case?
> By any chance we are running into below one? Can you check your gcc version?
> 
> https://lore.kernel.org/all/CAABZP2z=xu+07-y5fqFLidZz1VpSgrSwXa1mFHPb=b3ezr3...@mail.gmail.com/
> 
> Maybe CONFIG_DEBUG_PREEMPT worth a try to if it shows up anything.

Does not change anything, the kernel crashes all the same.

Nonetheless, I noticed that a tool that does some BPF tracing is
running, and without it the problem is not reproducible.

Thanks

Michal

Reply via email to