[PATCH] md: take reconfig_mutex when enabling PPL on an inactive array
From: Nguyen Ngoc Thang
Date: Thu Sep 24 2026 - 14:01:13 EST
consistency_policy_store() sets MD_HAS_PPL for an external-metadata array
that is not running yet, without holding reconfig_mutex. But ->pers is
only assigned after pers->run() returns, so "not running" also covers an
md_run() that is in progress under the lock.
If the write lands inside raid5_run() after log_init() has decided not to
set up PPL, or after a failed ppl_init_log() has cleared the flag, the
flag ends up set with conf->log_private still NULL. The next idle pass
of raid5d() then does:
raid5d() -> handle_active_stripes() -> log_flush_stripe_to_raid()
-> ppl_write_stripe_run()
and dereferences the NULL ppl_conf:
Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002
KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017]
RIP: 0010:ppl_write_stripe_run+0x60/0x1320 drivers/md/raid5-ppl.c:538
Take reconfig_mutex around the update and refuse it with -EBUSY once the
array is running, in which case ->change_consistency_policy() is the
right interface.
Fixes: 664aed04446c ("md: add sysfs entries for PPL")
Reported-by: syzbot+75d7e96ad03ac2dbbd9f@xxxxxxxxxxxxxxxxxxxxxxxxx
Closes: https://syzkaller.appspot.com/bug?extid=75d7e96ad03ac2dbbd9f
Signed-off-by: Nguyen Ngoc Thang <ngocthang2710.1999@xxxxxxxxx>
---
Notes:
This differs from the NULL check in ppl_write_stripe_run() that is being
patch-tested on the syzbot page: that hides the symptom but leaves
MD_HAS_PPL inconsistent with conf->log_private for the other PPL entry
points. md_attr_store() has not taken reconfig_mutex for every store since
6791875e2e53, so the store has to take it itself.
Testing (QEMU, KASAN, x86_64): one loop assembles raid5/external/ram0 while
a second loop keeps writing "ppl" to consistency_policy. Without a delay the
window is only a few microseconds, so I used a debug-only patch in
raid5_run() (clear MD_HAS_PPL and mdelay(50) before log_init()) that is not
part of this submission.
- without the fix: the syzbot oops on the first assembly,
ppl_write_stripe_run+0x5d in md0_raid5;
- with the fix and the same delay: 100/100 assemblies, no KASAN report.
The unmodified syz repro only ended with -ENOSPC ("PPL space too small") in
my setup.
drivers/md/md.c | 10 +++++++++-
1 file changed, 9 insertions(+), 1 deletion(-)
diff --git a/drivers/md/md.c b/drivers/md/md.c
index 680b34a63cb3..5b2add94631b 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -5884,7 +5884,15 @@ consistency_policy_store(struct mddev *mddev, const char *buf, size_t len)
else
err = -EBUSY;
} else if (mddev->external && strncmp(buf, "ppl", 3) == 0) {
- set_bit(MD_HAS_PPL, &mddev->flags);
+ /* md_run() holds the lock until ->pers is set */
+ err = mddev_lock(mddev);
+ if (err)
+ return err;
+ if (mddev->pers)
+ err = -EBUSY;
+ else
+ set_bit(MD_HAS_PPL, &mddev->flags);
+ mddev_unlock(mddev);
} else {
err = -EINVAL;
}
--
2.43.0