Re: [PATCH v4 1/6] rcu: Make call_rcu() safe to call from any context

From: Sebastian Andrzej Siewior

Date: Wed Aug 26 2026 - 10:47:10 EST


On 2026-08-10 05:27:50 [-0700], Puranjay Mohan wrote:
> --- a/kernel/rcu/rcu.h
> +++ b/kernel/rcu/rcu.h
> @@ -572,6 +572,17 @@ static inline void tasks_cblist_init_generic(void) { }
> #define RCU_SCHEDULER_INIT 1
> #define RCU_SCHEDULER_RUNNING 2
>
> +/*
> + * Defer whenever interrupts are disabled, since a callback-list operation may
> + * be in flight on this CPU. Not before the scheduler is up: irq_work is not
> + * usable that early, and rcu_init() itself calls call_rcu().
> + */
> +static inline bool should_rcu_defer(void)
> +{
> + return IS_ENABLED(CONFIG_RCU_DEFER) && irqs_disabled() &&
> + rcu_scheduler_active != RCU_SCHEDULER_INACTIVE;
> +}

Why does the description say that the defer part is for usage from NMI
and the test here has irqs_disabled() instead of in_nmi()?

> +
> enum rcutorture_type {
> RCU_FLAVOR,
> RCU_TASKS_FLAVOR,
> diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
> index 21b6ce1dffb63..3bf3a250f9de8 100644
> --- a/kernel/rcu/tree.c
> +++ b/kernel/rcu/tree.c
> @@ -3206,6 +3204,103 @@ __call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
> local_irq_restore(flags);
> }
>
> +/*
> + * Re-issue deferred callbacks straight to the enqueue so they cannot defer
> + * again. ->defer_lock serializes the drainers: this CPU's irq_work,
> + * rcu_defer_flush() and rcutree_migrate_callbacks().
> + */
> +static void __rcu_defer_drain(struct rcu_data *rdp)
> +{
> + struct llist_node *node, *next;
> + unsigned long flags;
> +
> + if (!IS_ENABLED(CONFIG_RCU_DEFER))
> + return;
> +
> + raw_spin_lock_irqsave(&rdp->defer_lock, flags);
> + llist_for_each_safe(node, next, llist_del_all(&rdp->defer_head)) {
> + struct rcu_head *head = (struct rcu_head *)node;

Why do you need the lock. This is still not clear to me despite the
comment. You can do llist_del_all() towards another list and then feed
it into rcu_do_enqueue() one by one. And you use the LAZY part.

> +
> + /* Bounds a node self-linked by a double call_rcu(). */
> + head->next = NULL;
> + rcu_do_enqueue(head, head->func, false);
> + }
> + raw_spin_unlock_irqrestore(&rdp->defer_lock, flags);
> +}

> @@ -4231,6 +4330,9 @@ rcu_boot_init_percpu_data(int cpu)
> rdp->rcu_onl_gp_state = RCU_GP_CLEANED;
> rdp->last_sched_clock = jiffies;
> rdp->cpu = cpu;
> + init_llist_head(&rdp->defer_head);
> + raw_spin_lock_init(&rdp->defer_lock);
> + rdp->defer_work = IRQ_WORK_INIT_HARD(rcu_defer_drain);

Why is this IRQ_WORK_INIT_HARD() instead, say, IRQ_WORK_INIT_LAZY()? Is
there a requirement that the RCU callback needs to complete asap and not
be delayed to the next tick? This would give kind of the LAZY part.

> rcu_boot_init_nocb_percpu_data(rdp);
> }
>

Sebastian