Re: [PATCH] mm/memory-tiers: cache top tier nodes

From: Joshua Hahn

Date: Tue Jul 14 2026 - 15:18:31 EST


On Sun, 12 Jul 2026 14:13:24 +0530 Aneesh Kumar K.V <aneesh.kumar@xxxxxxxxxx> wrote:

> Joshua Hahn <joshua.hahnjy@xxxxxxxxx> writes:
>
> > node_is_toptier() is called in a few hot paths: task_numa_fault(),
> > should_numa_migrate_memory(), folio_migrate_flags(), etc. Each call
> > takes an RCU read section and performs the tier distance check again.
> >
> > Tieredness for all nodes only changes on memory node hotplugs. Instead
> > of recomputing toptier nodes inside each of the hot paths above, compute
> > it only on memory node hotplug events and cache the results.
> >
> > Signed-off-by: Joshua Hahn <joshua.hahnjy@xxxxxxxxx>
> > ---
> > mm/memory-tiers.c | 36 ++++++++++--------------------------
> > 1 file changed, 10 insertions(+), 26 deletions(-)
> >
> > diff --git a/mm/memory-tiers.c b/mm/memory-tiers.c
> > index 54851d8a195b0..2e6e02ec1fce4 100644
> > --- a/mm/memory-tiers.c
> > +++ b/mm/memory-tiers.c
> > @@ -71,6 +71,7 @@ bool folio_use_access_time(struct folio *folio)
> >
> > #ifdef CONFIG_NUMA_MIGRATION
> > static int top_tier_adistance;
> > +static nodemask_t toptier_nodes __read_mostly = NODE_MASK_ALL;
> > /*
> > * node_demotion[] examples:
> > *
> > @@ -276,27 +277,7 @@ static struct memory_tier *__node_get_memory_tier(int node)
> > #ifdef CONFIG_NUMA_MIGRATION
> > bool node_is_toptier(int node)
> > {
> > - bool toptier;
> > - pg_data_t *pgdat;
> > - struct memory_tier *memtier;
> > -
> > - pgdat = NODE_DATA(node);
> > - if (!pgdat)
> > - return false;
> > -
> > - rcu_read_lock();
> > - memtier = rcu_dereference(pgdat->memtier);
> > - if (!memtier) {
> > - toptier = true;
> > - goto out;
> > - }
> > - if (memtier->adistance_start <= top_tier_adistance)
> > - toptier = true;
> > - else
> > - toptier = false;
> > -out:
> > - rcu_read_unlock();
> > - return toptier;
> > + return node_isset(node, toptier_nodes);
> > }
> >
> > void node_get_allowed_targets(pg_data_t *pgdat, nodemask_t *targets)
> > @@ -497,19 +478,22 @@ static void establish_demotion_targets(void)
> > }
> > }
> > /*
> > - * Now build the lower_tier mask for each node collecting node mask from
> > - * all memory tier below it. This allows us to fallback demotion page
> > - * allocation to a set of nodes that is closer the above selected
> > - * preferred node.
> > + * A node stays toptier unless it belongs to a tier below
> > + * top_tier_adistance, while each tier's lower_tier_mask collects the
> > + * nodes of every tier below it so demotion page allocation can fall
> > + * back to nodes closer to the selected preferred node.
> > */
> > + toptier_nodes = node_states[N_MEMORY];

Hello Aneesh, thank you very much for your review!

> Won't this result in a larger window where every node_is_toptier() call
> can return an incorrect result? This change would make every memory node
> a top-tier node until the loop below clears them from the toptier_nodes
> nodemask.

Yes, definitely. I think I should have instead constructed a local nodemask
and then swapped it out at the end, I overlooked this part.

> I assume an incorrect result from node_is_toptier() is not too bad, but
> is the performance impact large enough to make this change?

Yes, but I think it's still important to keep it correct.
On the notion of performance impact, it's not too big in the upstream
kernel, but I'm working on a series that adds more hotpaths that
frequently checks whether a folio belongs to a toptier node or not.
You can read more about it here [1].

Thank you again for your review, I'll send up a fixed version!
Joshua

[1] https://lore.kernel.org/all/20260423203445.2914963-1-joshua.hahnjy@xxxxxxxxx/