On Tue, Aug 11, 2026 at 04:17:41AM -0700, Breno Leitao wrote:
> On Mon, Aug 10, 2026 at 07:32:41PM +0300, Mike Rapoport wrote:
> > Hi Breno,
> > 
> > On Mon, Aug 10, 2026 at 06:32:04AM -0700, Breno Leitao wrote:
> > > Memory failures (such as unrecoverable ECCs errors) are getting more and
> > > more common. The kernel knows how to handle it while running, marking it
> > > as poisoned (and SIGBUS user tasks).
> > > 
> > > Poisoned memory is removed from the buddy allocator, but, not from
> > > other places. A current problem is that kexec will load new kernel
> > > on top of a bad/poisoned memory, which is undesirable.
> > > 
> > > If the next kernel's image, initrd or purgatory lands on poisoned frame,
> > > the relocation copy writes to the bad memory and the machine checks
> > 
> > What does the machine check here? ;-)
> 
> Not sure I got your question right. Did you mean:

I meant that "and the machine checks during the kexec" reads as machine
checks for something and that something is missing.

"machine check exceptions" would have been clearer :)
 
> 1) that there is no machine check exception when *writing* to poisoned
>    memory, or
> 
> 2) just that "the machine checks" is a lousy way to write it?
> 
> For 1) I think you are right, and I had not thought it through. The MCE
> (or a recurrent multi-bit ECC) would come from consuming the error, so
> a load or an instruction fetch, and a store may well pass silently and
> leave the poison sitting there.
> 
> The read back is what gets us, though. 
> 
> So the sentence should hang on the read, not on the copy. Would
> something like makes more sense?
> 
>         If the next kernel's image, initrd or purgatory lands on a
>         poisoned frame, the relocation copy puts them on memory that is
>         known bad.
> 
>         The error is consumed on the first read back, whether
>         that is purgatory checksumming the segments or the new kernel
>         running from them, and that is what we want to avoid.

I wouldn't overload the sentence, just 

        The error happens on the first from a bad page and that's what we
        want to avoid.

looks enough to me.
 
> > > Skip hardware-poisoned frames when placing segments: check them in the
> > > kexec_file hole finder so it lays the next kernel down on good memory,
> > > and reject a poisoned destination in sanity_check_segment_list() for
> > > the kexec_load path, which cannot relocate.
> > > 
> > > The two hole finders walk in opposite directions, so each asks for the
> > > end of the poison it has to clear: the top-down walk for the first
> > > poisoned page in the window, the bottom-up walk for the last. A poisoned
> > > hugetlb folio counts in full, as hugetlb keeps the flag on the folio and
> > > the poisoned subpages on its raw hwpoison list.
> > 
> > I had hard time parsing these two paragraphs. Can you please add more human
> > touch to them?
> 
> Sure, but that would cost more. :-) 
> 
> What about something like:
> 
>       Skip hardware-poisoned frames that were detected by machine

                                                          ^ memory

>       failure subssytem earlier when placing kexec segments.

               ^ subsystem
> 
>       To do so, add a helper that reports the first or the last poisoned page
>       in a range: memory is walked top-down by locate_mem_hole_top_down() and
>       bottom-up by locate_mem_hole_bottom_up(), so each direction needs a
>       different answer to jump clear of the poison.

                          ^ stay

>       kexec_load() gets its destinations from userspace and cannot move them,
>       so there sanity_check_segment_list() just rejects
>       a a segument/memory block that happens to have a posioned page.

        ^ single a 
> 
>       is_page_hwpoison() also covers hugetlb, where the flag sits on the folio
>       and the bad subpages on its raw hwpoison list, so a poisoned hugetlb
>       folio is skipped as a whole.

I don't think we care here about the list of bad subpages:

        is_page_hwpoison() also covers hugetlb, so a poisoned hugetlb
        folio is skipped as a whole.

> 
> > > +         poison = range_first_hwpoison(temp_start, kbuf->memsz);
> > > +         if (poison != PHYS_ADDR_MAX) {
> > > +                 /* we hit a poisoned page */
> > > +                 if (poison < kbuf->memsz)
> > > +                         return 0;
> > 
> > Won't we break out on the next iteration boundaries check? I.e.
> > 
> >             if (temp_start < start || temp_start < kbuf->buf_min)
> >                     return 0;
> 
> Kind-of.  Sashiko keeps raising this underflow in the function, on every
> revision since v2.
> 
> It dismisses it on this hunk because of the check, but reports it as
> a real one on the two "temp_start = temp_start - PAGE_SIZE" paths above,
> which do the same subtraction with nothing guarding them. 
> 
> Happy to remove it from here and send that as a separate patch.

Let's make it a separate patch please and drop the if (poison <
kbuf->memsz) here.
 
> > > + for (pfn = PHYS_PFN(start); pfn <= end_pfn; pfn++) {
> > > +
> > > +         cond_resched();
> > 
> > cond_resched() for every pfn is too much, isn't it?
> 
> It is what the other pfn walkers do: the kpageflags read loop in
> fs/proc/page.c and read_page_owner() in mm/page_owner.c both call it
> once per pfn. 

I think it depends on the pfn walker, some of them cond_resched() once per
"block"
 
> But I honestly don't have a strong opinion here, though, happy to batch
> it if you prefer. Would this one look better?:
> 
>       if (!(pfn % MAX_ORDER_NR_PAGES))
>               cond_resched();

Can't say I know the magic number here, but I think it's better to batch.
We had a related discussion with Muchun a short while ago:

https://lore.kernel.org/all/[email protected]/
 
> Thanks for the review,
> --breno

-- 
Sincerely yours,
Mike.

Reply via email to