Re: [PATCH 0/2] printk/stop_machine: Defer legacy console flushes while a CPU runs a stopper callback
From: John Ogness
Date: Fri Aug 28 2026 - 04:56:17 EST
Hi Aditya,
On 2026-08-27, Aditya Chillara <aditya.chillara@xxxxxxxxxxxxxxxx> wrote:
> A device using a legacy UART console (console=ttyMSM0,115200n8) hit a
> watchdog bark/bite about 40 seconds after boot.
>
> stop_machine() (used here for kprobe text patching) stops every CPU by
> running multi_cpu_stop() on each of them, through the per-CPU
> "migration/%u" threads. These threads run at a higher priority than the
> msm_watchdog thread. At bite time, all eight CPUs were still spinning in
> multi_cpu_stop()'s MULTI_STOP_PREPARE state, where interrupts are left
> enabled.
>
> Heavy SELinux denial logging had built up a large backlog on the
> console. One CPU took an interrupt while spinning in MULTI_STOP_PREPARE.
> Handling it eventually led to a printk(), and because the console was a
> legacy console, that printk() synchronously drained the whole backlog
> over the slow UART. While the drain was still running, the watchdog bark
> interrupt hit the same CPU, found no recent pet, and escalated to a
> bite.
>
> The captured stack for that CPU, innermost frame first:
>
> qcom_soc_set_wdt_bite
> qcom_wdt_bark_handler
> __handle_irq_event_percpu
> handle_irq_event
> handle_fasteoi_irq
> generic_handle_domain_irq
> gic_handle_irq
> do_interrupt_handler
> el1_interrupt
> el1h_64_irq_handler
> el1h_64_irq
> console_flush_all
> console_unlock
> vprintk_emit
> dev_vprintk_emit
> dev_printk_emit
> __dev_printk
> _dev_err
> btspi_sleep_timeout_handler
> call_timer_fn
> __run_timer_base
> run_timer_softirq
> handle_softirqs
> __do_softirq
> ____do_softirq
> call_on_irq_stack
> do_softirq_own_stack
> __irq_exit_rcu
> irq_exit_rcu
> el1_interrupt
> el1h_64_irq_handler
> el1h_64_irq
> multi_cpu_stop
> cpu_stopper_thread
> smpboot_thread_fn
> kthread
> ret_from_fork
>
> Every other CPU stayed parked in the rendezvous the whole time, since
> their stopper threads outrank msm_watchdog. Nothing could pet the
> watchdog until the drain finished.
>
> This was observed through multi_cpu_stop(), but the hazard is not
> specific to it. Every cpu stopper callback runs in stop_sched_class,
> above msm_watchdog and every other thread on the CPU, so a slow flush
> from any of them (including single-CPU callbacks such as the migration
> and task-migration stoppers) can starve the watchdog just as well. The
> fix therefore covers all stopper callbacks, not only multi_cpu_stop().
>
> Fix this by having the cpu stopper mark the CPU active while a callback
> runs, and having printk use that marker to defer legacy console flushes
> until the callback returns:
>
> 1/2 stop_machine: Track when a CPU executes a stopper callback
>
> Add a per-CPU flag, set in the stopper dispatch path around the
> callback, and an in_cpu_stop() accessor.
>
> 2/2 printk: Defer legacy console flushes while a CPU runs a stopper callback
>
> Route legacy console output through the offload path instead of
> flushing it directly while a CPU is inside a stopper callback, and
> flush it once the callback returns. Emergency and panic output is
> unaffected.
>
> Reproduced and verified with an out-of-tree test module that triggers
> stop_machine() with a queued console backlog and a printk() inside the
> rendezvous, paired with a kprobe-based script that flags any console
> flush happening while a CPU is inside a stopper callback.
Generally speaking, we are not taking the whack-a-mole approach to
workaround all the known legacy console problems (there are a lot of
them!). However, if there are problems that occur during normal usage
(as opposed to crafted tests), then we can insert workarounds.
For workarounds of known legacy console problems we have the deferred
enter/exit functions. These only affect legacy consoles and literally
exist for these purposes. I would expect the following patch would also
solve your problem.
John Ogness
diff --git a/kernel/stop_machine.c b/kernel/stop_machine.c
index d085ba1f4b44e..31f7af41249f1 100644
--- a/kernel/stop_machine.c
+++ b/kernel/stop_machine.c
@@ -507,7 +507,9 @@ static void cpu_stopper_thread(unsigned int cpu)
stopper->caller = work->caller;
stopper->fn = fn;
preempt_count_inc();
+ printk_deferred_enter();
ret = fn(arg);
+ printk_deferred_exit();
if (done) {
if (ret)
done->ret = ret;