Re: [PATCH] cgroup/cpuset: Invalidate remote partition on housekeeping conflict

From: Waiman Long

Date: Fri Sep 25 2026 - 14:15:34 EST


On 9/23/26 10:03 PM, Ridong Chen wrote:


On 9/23/2026 6:36 PM, Guopeng Zhang wrote:
From: Guopeng Zhang <zhangguopeng@xxxxxxxxxx>

Widening an ancestor's exclusive CPU mask can add a boot-isolated CPU
to a valid remote partition root without touching the partition's own
control files. The partition then load balances that CPU, silently
defeating isolcpus=domain for it.

This can be reproduced on a 32-CPU system booted with
isolcpus=domain,4:

     cd /sys/fs/cgroup
     echo +cpuset > cgroup.subtree_control
     mkdir -p A/B
     echo +cpuset > A/cgroup.subtree_control
     echo 2-4 > A/cpuset.cpus
     echo 2-3 > A/cpuset.cpus.exclusive
     echo 2-4 > A/B/cpuset.cpus
     echo 2-4 > A/B/cpuset.cpus.exclusive
     echo root > A/B/cpuset.cpus.partition
     cat A/B/cpuset.cpus.effective               # 2-3
     echo 2-4 > A/cpuset.cpus.exclusive
     cat A/B/cpuset.cpus.partition               # root
     cat A/B/cpuset.cpus.effective               # 2-4

The last write returns 0 and leaves the hierarchy in this state:

     root (cpuset.cpus.effective=0-1,5-31)
     |
     \-- A (member):           cpuset.cpus=2-4
         |                     cpuset.cpus.exclusive=2-4
         \-- B (root, remote): cpuset.cpus=2-4
                               cpuset.cpus.effective=2-4

B is a remote partition: it takes its CPUs directly from the root
cpuset, and A only passes its exclusive list down. Before the last
write, that list is 2-3, so B holds 2-3 and CPU 4 stays in the
root cpuset as a boot-isolated CPU. The write widens A's exclusive
list to 2-4, which additionally grants CPU 4 to B. Nothing rejects
the grant: B remains a valid root partition, and CPU 4 is still
listed in cpuset.cpus.isolated while sitting in a load-balanced
partition.

remote_partition_enable() and validate_partition() already call
prstate_housekeeping_conflict() before granting CPUs. The ancestor
update path in remote_cpus_update() does not. Add the missing check so
that the existing prs_err path invalidates the remote partition instead
of adding the boot-isolated CPU.

Fixes: f62a5d39368e ("cgroup/cpuset: Remove remote_partition_check() & make update_cpumasks_hier() handle remote partition")
Signed-off-by: Guopeng Zhang <zhangguopeng@xxxxxxxxxx>
---
  kernel/cgroup/cpuset.c | 2 ++
  1 file changed, 2 insertions(+)

diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 753aa65afcd7..362e5b5dccaa 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -1699,6 +1699,8 @@ static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
          else if ((prs == PRS_ISOLATED) &&
               !isolated_cpus_can_update(tmp->addmask, tmp->delmask))
              WRITE_ONCE(cs->prs_err, PERR_HKEEPING);
+        else if (prstate_housekeeping_conflict(prs, tmp->addmask))
+            WRITE_ONCE(cs->prs_err, PERR_HKEEPING);
          if (cs->prs_err)
              goto invalidate;
      }

Hi Guopeng,

Thank you for your test and patch.

I noticed that this statement existed in remote_partition_enable but was missing from remote_cpus_update. There may be other state leaks here as well.

In case this happens again, I would like to suggest adding 'validate_remote_partition' as I tried to do before [1]. Could you put some effort into this?

[1] https://lore.kernel.org/cgroups/20251225123058.231765-19-chenridong@xxxxxxxxxxxxxxx/

I agreed that we should consolidate similar code into common helper that will make the code to easier to maintain and debug and avoid this kind of missing check problem.

Cheers,
Longman