Re: [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails

From: Tejun Heo

Date: Mon Sep 28 2026 - 22:07:31 EST


Hello, Waiman.

The following is a Claude-generated review.

On Mon, Sep 28, 2026 at 07:51:59PM -0400, Waiman Long wrote:
> Before commit 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full
> don't leave any housekeeping"), the partition state can be freely switched
> from "root" to "isolated" and vice versa. After that commit, the switch
> from "root" to "isolated" can fail if it exhausts all the housekeeping
> CPUs. Later on, the switch from "isolated" to "root" can also fail if
> some of the partition CPUs are boot-time isolated by "isolcpus".

The isolated -> root failure came from b1034a690129 ("cgroup/cpuset:
Ensure domain isolated CPUs stay in root or isolated partition"). Maybe
add a Fixes: tag for it too?

> The partition is made invalid when the switch fails. However, the
> remote_partition flag for a remote partition can remain set and the CPUs
> from the invalidated partition aren't cleared from subpartitions_cpus.
> Fix this by properly disable the partition in this case.

The reproducer below is a local partition and the fix covers that case
too. The CPUs weren't given back to the parent, which is what leaves them
in subpartitions_cpus and isolated_cpus in the example. Maybe describe
both? Also, "properly disable" -> "disabling".

> In the case of remote_partition flag, it should be cleared for a
> invalidated remote partition. To be safe, the reset_partition_data() is
> now enhanced to always clear the remote_partition flag. So there is no
> need to explicitly clear remote_partition in remote_partition_disable().

After the update_prstate() change, every path that invalidates a remote
partition goes through remote_partition_disable(), so the other
reset_partition_data() callers never see the flag set. If one did,
clearing only the flag would leave its CPUs in subpartitions_cpus and turn
the WARN_ON_ONCE() in partition_xcpus_del() into a silent leak. Maybe drop
this part?

> On a x86 test system with boot option "isolcpus=10 cgroup_debug" set
> and more than 16 cores, the following commands was executed after boot.

"a x86" -> "an x86", "commands was" -> "commands were".

> @@ -2946,27 +2947,32 @@ static int update_prstate(struct cpuset *cs, int new_prs)
> */
> if (((new_prs == PRS_ISOLATED) &&
> !isolated_cpus_can_update(cs->effective_xcpus, NULL)) ||
> - prstate_housekeeping_conflict(new_prs, cs->effective_xcpus))
> + prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) {
> err = PERR_HKEEPING;
> - else
> + disable_partition = true;

If a root -> isolated switch fails isolated_cpus_can_update() under an
isolated parent, partcmd_disable hands the CPUs back to the parent and
partition_xcpus_del() adds them to isolated_cpus, which is the state the
check just rejected. Switching to member ends up in the same place, so
this may be fine as is.

> + }
> +out:
> + if (disable_partition) {

The early goto out paths never need the disable. Maybe put this block
before out: instead?

Also, update_cpumasks_hier() below gets force only when switching to
member, to update effective_xcpus. Now that the failure path disables the
partition too, should it pass disable_partition?

Separately, a partition invalidated with PERR_HKEEPING can become valid
again through partcmd_update without newmask (hotplug, or
update_cpumasks_hier() from an ancestor), which doesn't check
housekeeping. A failed member -> root enable has the same problem, so it
isn't from this patch.

Thanks.

--
tejun