Re: [PATCH 1/1] mm/memory_hotplug: fix missing rollback in __add_pages()
From: Lance Yang
Date: Thu Oct 01 2026 - 01:58:10 EST
On 2026/10/1 00:46, David Hildenbrand (Arm) wrote:
On 9/30/26 18:04, Lance Yang wrote:
__add_pages() returns on a sparse_add_section() failure without removing
the sections already added in the same request.
For memremap_pages(), the failed range is not counted in pgmap->nr_range,
so memunmap_pages() skips it. The sections already added in that range
retain their vmemmap mappings and subsection bits. Retrying a
section-aligned range can then fail with -EEXIST.
Save the initial PFN and remove [start_pfn, pfn) on failure. For a vmemmap
population failure, section_activate() already cleans up the current
section, so the rollback excludes it. If the first section fails,
__remove_pages() receives an empty range and does nothing.
Link: https://lore.kernel.org/all/BAD58999-1EDD-4A37-ABA0-DB1BD8AB3453@xxxxxxxxx/
Suggested-by: Muchun Song <muchun.song@xxxxxxxxx>
Signed-off-by: Lance Yang <lance.yang@xxxxxxxxx>
---
No Fixes tag, as I couldn't identify the commit that introduced this issue.
mm/memory_hotplug.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
index 796af1028ee2..ca4656698148 100644
--- a/mm/memory_hotplug.c
+++ b/mm/memory_hotplug.c
@@ -380,6 +380,7 @@ EXPORT_SYMBOL_GPL(pfn_to_online_page);
int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
struct mhp_params *params)
{
+ const unsigned long start_pfn = pfn;
const unsigned long end_pfn = pfn + nr_pages;
unsigned long cur_nr_pages;
int err;
@@ -413,8 +414,11 @@ int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
SECTION_ALIGN_UP(pfn + 1) - pfn);
err = sparse_add_section(nid, pfn, cur_nr_pages, altmap,
params->pgmap);
- if (err)
+ if (err) {
+ __remove_pages(start_pfn, pfn - start_pfn, altmap,
+ params->pgmap);
break;
+ }
cond_resched();
}
vmemmap_populate_print_last();
Makes sense and LGTM.
Do we have a Fixes: tag? It probably dates back quite a while ... not sure about
stable, we never saw this in practice. But if it's easy, we should just do it?
(not sure if we ever had __remove_pages be limited to hotunplug support)
Hmm ... I'd leave stable out for now. David, could you add the Fixes tag Muchun
suggested?
Fixes: ba72b4c8cf60 ("mm/sparsemem: support sub-section hotplug")
I'm planning on picking this up and sending it for the next merge window (so not
as a hotfix).
Thank you ;)
Looking at this ...
x86 does not really expect add_pages to fail:
ret = __add_pages(nid, start_pfn, nr_pages, params);
WARN_ON_ONCE(ret);
That's probably something to clean up as well?
That would be a separate fix, I guess :) I'll have a look at the other
architectures while I'm at it :)
Cheers, Lance