[PATCH-next 1/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only
From: Waiman Long
Date: Sun Sep 27 2026 - 12:37:03 EST
Commit f82f80426f7a ("sched/deadline: Ensure that updates to exclusive
cpusets don't break AC") adds a check in validate_change() to make
sure that there is enough bandwidth for SCHED_DEADLINE tasks if we
shrink a v1 exclusive cpuset that has CS_CPU_EXCLUSIVE flag set.
With the introduction of cpuset partition in cgroup v2, we keep setting
the CS_CPU_EXCLUSIVE flag for a partition root so that the SCHED_DEADLINE
check will continue to work as intended. However it turns out that the
current code isn't perfect and there are cases where a cpuset isn't
a valid partition root, but the exclusive flag is still incorrectly
set. This can leads to SCHED_DEADLINE check being incorrectly triggered
when there are deadline tasks in the system. This can result in unexpected
-EBUSY failure when making changes to cpuset control files. Fix that
by checking for a valid partition root in the case of v2 and break out
the v1 specific check back into cpuset1_validate_change(). It is far
easier and less cumbersome than to make sure that the exclusive flag
is only set for valid partition roots.
The is_in_v2_mode() check guarding the call to cpuset1_validate_change()
is also moved to the appropriate place inside cpuset1_validate_change()
and a cpuset_v2() is now being used as a guard which should be more
accurate for the a v1 system with v2 mode enabled.
Even though commit a86ce68078b2 ("cgroup/cpuset: Extract out
CS_CPU_EXCLUSIVE & CS_SCHED_LOAD_BALANCE handling") is marked as a
commit to be fixed, the problem may exist before that.
Fixes: a86ce68078b2 ("cgroup/cpuset: Extract out CS_CPU_EXCLUSIVE & CS_SCHED_LOAD_BALANCE handling")
Signed-off-by: Waiman Long <longman@xxxxxxxxxx>
---
kernel/cgroup/cpuset-v1.c | 20 ++++++++++++++++++--
kernel/cgroup/cpuset.c | 17 +++++------------
2 files changed, 23 insertions(+), 14 deletions(-)
diff --git a/kernel/cgroup/cpuset-v1.c b/kernel/cgroup/cpuset-v1.c
index 562ad35f00d0..c03ae8aac03a 100644
--- a/kernel/cgroup/cpuset-v1.c
+++ b/kernel/cgroup/cpuset-v1.c
@@ -61,6 +61,11 @@ struct cpuset_remove_tasks_struct {
#define FM_MAXCNT 1000000 /* limit cnt to avoid overflow */
#define FM_SCALE 1000 /* faux fixed point scale */
+static inline bool is_in_v2_mode(void)
+{
+ return cpuset_cgrp_subsys.root->flags & CGRP_ROOT_CPUSET_V2_MODE;
+}
+
/* Initialize a frequency meter */
static void fmeter_init(struct fmeter *fmp)
{
@@ -357,10 +362,21 @@ int cpuset1_validate_change(struct cpuset *cur, struct cpuset *trial)
if (!is_cpuset_subset(c, trial))
goto out;
- /* On legacy hierarchy, we must be a subset of our parent cpuset. */
+ /*
+ * We can't shrink if we won't have enough room for SCHED_DEADLINE
+ * tasks in a scheduling partition.
+ */
+ if (is_cpu_exclusive(cur) && is_sched_load_balance(cur) &&
+ !cpuset_cpumask_can_shrink(cur->cpus_allowed, trial->cpus_allowed))
+ goto out;
+
+ /*
+ * On legacy hierarchy with v2 mode off, we must be a subset of our
+ * parent cpuset.
+ */
ret = -EACCES;
par = parent_cs(cur);
- if (par && !is_cpuset_subset(trial, par))
+ if (par && !is_in_v2_mode() && !is_cpuset_subset(trial, par))
goto out;
/*
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 753aa65afcd7..5639c486c967 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -752,7 +752,7 @@ static int validate_change(struct cpuset *cur, struct cpuset *trial)
rcu_read_lock();
- if (!is_in_v2_mode())
+ if (!cpuset_v2())
ret = cpuset1_validate_change(cur, trial);
if (ret)
goto out;
@@ -765,23 +765,16 @@ static int validate_change(struct cpuset *cur, struct cpuset *trial)
/*
* We can't shrink if we won't have enough room for SCHED_DEADLINE
- * tasks. This check is not done when scheduling is disabled as the
- * users should know what they are doing.
- *
- * For v1, effective_cpus == cpus_allowed & user_xcpus() returns
- * cpus_allowed.
- *
- * For v2, is_cpu_exclusive() & is_sched_load_balance() are true only
- * for non-isolated partition root. At this point, the target
- * effective_cpus isn't computed yet. user_xcpus() is the best
- * approximation.
+ * tasks. This check is only done on non-isolated partition root.
+ * At this point, the target effective_cpus isn't computed yet.
+ * user_xcpus() is the best approximation.
*
* TBD: May need to precompute the real effective_cpus here in case
* incorrect scheduling of SCHED_DEADLINE tasks in a partition
* becomes an issue.
*/
ret = -EBUSY;
- if (is_cpu_exclusive(cur) && is_sched_load_balance(cur) &&
+ if (is_partition_valid(cur) && is_sched_load_balance(cur) &&
!cpuset_cpumask_can_shrink(cur->effective_cpus, user_xcpus(trial)))
goto out;
--
2.55.0