Confirming the upstream result with more data, and moving on to test
the current Ubuntu kernel

Following Mario's conclusion that this looks like a backport issue:
here is the accumulated
evidence on mainline, and a test of 7.0.0-29 that I had not run yet.

Mainline 7.2-rc5, with BOTH mitigations enabled and no workaround on
the command line:

    boot params      nmi_watchdog=1        (no retbleed=off, no
spec_rstack_overflow=off,
                                            no processor.max_cstate=1)
    retbleed         Mitigation: untrained return thunk; SMT enabled
with STIBP protection
    spec_rstack_overflow  Mitigation: Safe RET

    7 boots, 332 h accumulated (13.8 days)
    longest single boot 4d06h, which is 2.9x the upper end of the
original 12-36 h window
    current boot 3d05h and counting
    zero call traces, zero oops, zero panics

For contrast, the same machine on the Ubuntu kernels needed
retbleed=off to stay up, and the
crash reproduced within 12 to 36 h without it.

Hardware, unchanged from the original report:

    AMD Ryzen 7 4800U (family 23, model 96, stepping 1)
    microcode 0x860010d
    Beelink SER, BIOS SER_V1.14_P3C6M43_B_Link (2022-03-24)

I noticed 7.0.0-29 has been installed here since 4 August and was
never booted: every boot since
27 July has been on mainline. Since it postdates the 7.0.0-28 that
crashed, it may already carry
a fix, and testing it is the datapoint that is actually missing from
this report. Booting into it
now, with both mitigations enabled and no workaround, and will report
back either way.

kdump is armed (512 MB reserved, "ready to kdump"), so if it panics
there will be a vmcore with
the full stack rather than another description of a screen.

One methodological note that may be useful to whoever picks this up.
Searching the journal for
the crash does not work: a hard panic here kills the machine before
journald flushes, so after
the reboot there is simply no record. I verified this across 32
historical boots, and the only
"Call Trace" style matches on the crashing kernel turned out to be
userspace faults from an
unrelated application. The reliable signals are the vmcore and the
plain uptime reached.

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2160457

Title:
   Kernel panic in srso_safe_ret / kick_ilb during sched_tick on AMD
  Ryzen 7 4800U (Zen2)

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2160457/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to