Re: [RFC PATCH v2 02/23] sched/topology: Introduce a NUMA distance matrix with unique distance values
From: Peter Zijlstra
Date: Mon Aug 31 2026 - 07:55:53 EST
On Thu, Aug 27, 2026 at 08:27:55PM +0800, Jianyong Wu wrote:
> Builds a refined node distance matrix based on the raw NUMA distance matrix
> provided by BIOS. The refined matrix preserves the relative ordering of
> NUMA distances, while assigning distinct distance values to node pairs that
> originally shared identical distances within each matrix row. This matrix
> is exclusively used for cache-aware scheduling and has no impact on existing
> NUMA topology logic such as sched domain construction.
>
> For example, consider a system with 4 NUMA nodes. The raw BIOS-provided
> distance matrix may look like this:
>
> NODE0 NODE1 NODE2 NODE3
> NODE0 10 20 20 30
> NODE1 20 10 20 25
> NODE2 20 20 10 20
> NODE3 30 25 20 10
>
> Multiple duplicate distance values exist within each row. After the
> deduplication step, the refined distance matrix becomes:
>
> NODE0 NODE1 NODE2 NODE3
> NODE0 10 15 20 30
> NODE1 15 10 12 25
> NODE2 20 12 10 15
> NODE3 30 25 15 10
>
> All entries in each row are now unique, while adhering to two core principles:
> 1. The relative distance ordering from the original matrix is preserved.
> For instance, original distance(NODE0, NODE1) < distance(NODE0, NODE3),
> and this relative relationship is retained in the refined matrix as well.
> 2. The matrix remains symmetric across its main diagonal. Maintaining
> symmetry is critical to guarantee consistent pairwise node distances.
This example uses Node only, but the code in question is specifically
aimed at Cache granularity; might it be better to use a cache example?
A little something like so (I got tired of prompting Gemini to generate
more complicates / less broken examples)...
Pre:
Cache | C0 C1 | C2 C3 | C4 C5 | C6 C7
------+----------+----------+----------+---------
C0 | 10 10 | 20 20 | 20 20 | 20 20
C1 | 10 10 | 20 20 | 20 20 | 20 20
------+----------+----------+----------+---------
C2 | 20 20 | 10 10 | 20 20 | 20 20
C3 | 20 20 | 10 10 | 20 20 | 20 20
------+----------+----------+----------+---------
C4 | 20 20 | 20 20 | 10 10 | 20 20
C5 | 20 20 | 20 20 | 10 10 | 20 20
------+----------+----------+----------+---------
C6 | 20 20 | 20 20 | 20 20 | 10 10
C7 | 20 20 | 20 20 | 20 20 | 10 10
Post:
Cache | C0 C1 | C2 C3 | C4 C5 | C6 C7
------+----------+----------+----------+---------
C0 | 10 11 | 20 21 | 22 23 | 24 25
C1 | 11 10 | 21 20 | 23 22 | 25 24
------+----------+----------+----------+---------
C2 | 20 21 | 10 11 | 24 25 | 22 23
C3 | 21 20 | 11 10 | 25 24 | 23 22
------+----------+----------+----------+---------
C4 | 22 23 | 24 25 | 10 11 | 20 21
C5 | 23 22 | 25 24 | 11 10 | 21 20
------+----------+----------+----------+---------
C6 | 24 25 | 22 23 | 20 21 | 10 11
C7 | 25 24 | 23 22 | 21 20 | 11 10
> Each row of this refined NUMA distance matrix is sorted in ascending order to
> generate a unique per-node affinity sequence. This sequence will guide
> thread migration logic introduced in subsequent patches.
IIRC greedy has significant worse bounds than many other schemes. This
would result in more unique distances than strictly needed here, right?
Since this is all on slow paths anyway, does it make sense to pick a
slightly better algorithm in order to reduce this bound and get better
results?
Anyway, let me continue trying to dig through all this.