[PATCH] mm/slab: disallow kfree_rcu_sheaf() on PREEMPT_RT again
From: Vlastimil Babka (SUSE)
Date: Mon Aug 31 2026 - 15:50:03 EST
This partially reverts commit 2a8bb29ec9b2 ("mm/slab: allow
kfree_rcu_sheaf() on PREEMPT_RT"). It was based on an assumption that
local_trylock() is safe on PREEMPT_RT from any context.
However kvfree_rcu() is also called by set_cpus_allowed_force() with
task_struct::pi_lock acquired and there it's not safe, as syzbot
has reported.
For the immediate fix, skip kfree_rcu_sheaf() on PREEMPT_RT again from
kvfree_call_rcu(). In theory, kfree_rcu_nolock() would have the same
problem when called from under pi_lock on PREEMPT_RT but that can
be addressed if such a caller is proposed.
Add an explanation comment, courtesy of Sebastian.
Reported-by: syzbot+acf142088e0182172e58@xxxxxxxxxxxxxxxxxxxxxxxxx
Closes: https://syzkaller.appspot.com/bug?extid=acf142088e0182172e58
Reported-by: ThangNN99 <ngocthang2710.1999@xxxxxxxxx>
Fixes: 2a8bb29ec9b2 ("mm/slab: allow kfree_rcu_sheaf() on PREEMPT_RT")
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@xxxxxxxxxx>
---
Incidentally I have posted a RFC [1] that leads to replacing that
kfree_rcu() from set_cpus_allowed_force() but now after back from
vacation I need to check the feedback and based on this bug report I can
already see it makes the same bad assumption that trylock is fine.
[1] https://lore.kernel.org/all/20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@xxxxxxxxxx/
---
mm/slab_common.c | 18 ++++++++----------
mm/slub.c | 5 +++--
2 files changed, 11 insertions(+), 12 deletions(-)
diff --git a/mm/slab_common.c b/mm/slab_common.c
index b19ba1b31484..7223a7596dab 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -1667,14 +1667,6 @@ static bool kfree_rcu_sheaf(void *obj)
{
struct kmem_cache *s;
struct slab *slab;
- unsigned int free_flags = SLAB_FREE_DEFAULT;
-
- /*
- * It is not safe to spin on PREEMPT_RT because the kernel might be
- * holding a raw spinlock and slab acquires sleeping locks.
- */
- if (IS_ENABLED(CONFIG_PREEMPT_RT))
- free_flags = SLAB_FREE_NOLOCK;
if (is_vmalloc_addr(obj))
return false;
@@ -1685,7 +1677,7 @@ static bool kfree_rcu_sheaf(void *obj)
s = slab->slab_cache;
if (likely(!IS_ENABLED(CONFIG_NUMA) || slab_nid(slab) == numa_mem_id()))
- return __kfree_rcu_sheaf(s, obj, free_flags);
+ return __kfree_rcu_sheaf(s, obj, SLAB_FREE_DEFAULT);
return false;
}
@@ -2034,7 +2026,13 @@ void kvfree_call_rcu(struct kvfree_rcu_head *head, void *ptr)
if (!head)
might_sleep();
- if (kfree_rcu_sheaf(ptr))
+ /*
+ * kvfree_rcu() is called by set_cpus_allowed_force() with
+ * task_struct::pi_lock acquired. On PREEMPT_RT the local_trylock()
+ * usage below will acquire the waitlock which must be avoided.
+ * Therefore avoid it on PREEMPT_RT.
+ */
+ if (!IS_ENABLED(CONFIG_PREEMPT_RT) && kfree_rcu_sheaf(ptr))
return;
// Queue the object but don't yet schedule the batch.
diff --git a/mm/slub.c b/mm/slub.c
index f9b56cb439e7..7a7e906a0e44 100644
--- a/mm/slub.c
+++ b/mm/slub.c
@@ -6088,8 +6088,9 @@ static void rcu_free_sheaf(struct rcu_head *head)
/*
* kvfree_call_rcu() can be called while holding a raw_spinlock_t. Since
* __kfree_rcu_sheaf() may acquire a spinlock_t (sleeping lock on PREEMPT_RT),
- * this would violate lock nesting rules. Therefore, kvfree_call_rcu() avoids
- * this problem by passing SLAB_FREE_NOLOCK on PREEMPT_RT.
+ * this would violate lock nesting rules. Therefore, kfree_call_rcu_nolock()
+ * avoids this problem by passing SLAB_FREE_NOLOCK. kvfree_call_rcu() is
+ * bypassing the sheaves layer completely on PREEMPT_RT.
*
* However, lockdep still complains that it is invalid to acquire spinlock_t
* while holding raw_spinlock_t, even on !PREEMPT_RT where spinlock_t is a
---
base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
change-id: 20260831-b4-kfree_rcu_hotfix-1d3dd315bff0