Re: [PATCH v3 03/10] crash_dump: Disallow writing to dm-crypt configfs during kexec_file_load syscall
From: Coiby Xu
Date: Wed Aug 19 2026 - 06:12:20 EST
On Sun, Aug 09, 2026 at 05:26:54PM +0530, Sourabh Jain wrote:
Hello Coiby,
On 06/08/26 10:50, Coiby Xu wrote:
On Wed, Aug 05, 2026 at 04:00:51PM +0530, Sourabh Jain wrote:
On 29/07/26 09:06, Coiby Xu wrote:
If writing to the configfs group happens concurrently during
kexec_file_load syscall, it may lead to the following issues,
- buffer overflow if dm-crypt keys are added after allocation
- stale total_keys if dm-crypt keys are removed during iteration
- keys_header will not be freed if config/crash_dm_crypt_key/reuse is
set true
So hold config_keys_subsys.su_mutex for the entire sequence during the
kexec_file_load syscall to ensure a consistent snapshot.
Fixes: 479e58549b0f ("crash_dump: store dm crypt keys in kdump reserved memory")
Suggested-by: Sourabh Jain <sourabhjain@xxxxxxxxxxxxx>
Signed-off-by: Coiby Xu <coiby.xu@xxxxxxxxx>
---
kernel/crash_dump_dm_crypt.c | 23 +++++++++++++++++++++--
1 file changed, 21 insertions(+), 2 deletions(-)
diff --git a/kernel/crash_dump_dm_crypt.c b/kernel/crash_dump_dm_crypt.c
index 4335b6cb1fc4..d2e66c6fe6f3 100644
--- a/kernel/crash_dump_dm_crypt.c
+++ b/kernel/crash_dump_dm_crypt.c
@@ -293,6 +293,7 @@ static ssize_t config_keys_reuse_show(struct config_item *item, char *page)
static ssize_t config_keys_reuse_store(struct config_item *item,
const char *page, size_t count)
{
+ struct mutex *lock;
bool val;
int r;
@@ -302,8 +303,12 @@ static ssize_t config_keys_reuse_store(struct config_item *item,
return -EINVAL;
}
+ lock = &to_config_group(item)->cg_subsys->su_mutex;
+ mutex_lock(lock);
Is this lock only protecting against races between key reuse and kexec_file_load(),
The lock here is to protect against races between key reuse and
kexec_file_load.
or does it also handle the case where a new key is added during key reuse or
kexec_file_load() is running?
For the cases where a key is added/deleted, configfs will automatically
take care of them because it will acquire mutex lock automatically.
If it is only intended to protect key reuse versus kexec_file_load(), why can't we use
the kexec lock instead?
The reason I'm asking is that, in upcoming patches, the key reuse path accesses
kexec_crash_image properties and the crash reserved region directly. Doing so
without taking the kexec lock (using kexec_trylock()) could lead to race conditions.
After comparing the kexec lock approach with the configfs mutex lock
approach, I think the latter is a simpler solution because
1. the kexec lock is non-blocking and we have to repeatedly try until the
lock get acquired. So it means user space has to make changes as
well.
2. configfs already acquires the mutex lock automatically for
creating/deleting configfs items. So if we use configfs mutex lock,
it means one less place to use the lock.
In config_keys_reuse_store, kexec_crash_image will be checked before
accessing its properties and the crash reserved region. Can you
elaborate on what the race conditions are? Will acquiring the lock
before accessing kexec_crash_image properties and the crash reserved
region help protect against these races?
In theory, the kexec lock can be a more robust approach. But considering
only root can write to the crash dm-crypt keys configfs and load kdump
image, I'm not sure it's necessary to adopt a bit more complex solution.
The reuse function accesses the kexec crash image properties and crashkernel
memory without taking the kexec lock. This could lead to race conditions or
other problems.
Since we need to take the kexec lock anyway, my suggestion is: can we use the
kexec lock instead of cg_subsys->su_mutex if the purpose of the mutex is only
to synchronize reuse with kexec_file_load?
As we know, kexec_file_load already runs under the kexec lock. So using the
same lock should provide the required synchronization.
- Sourabh Jain
After 1) digging into the git history to learn more about what problems may
happen without ensuring serial access to the kexec crash image
properties and 2) noticing kexec-tools won't retry when kexec_file_load
syscall failed with -EBUSY due to kexec lock acquisition failure which
implies it's very unlikely to fail to get the kexec lock, I'm convinced
it's better to switch to kexec lock for the reuse function. Thanks for
the suggestion!
--
Best regards,
Coiby