Re: [PATCH v2] ACPI: APEI: Handle repeated SEA error storms
From: hejunhao
Date: Thu Jul 23 2026 - 04:33:08 EST
On 2026/7/9 1:14, Luck, Tony wrote:
> On Wed, May 27, 2026 at 04:27:07PM +0800, Junhao He wrote:
>> When hardware memory corruption occurs and a user process accesses the
>> corrupted page, the CPU triggers a Synchronous External Abort (SEA).
>> The kernel invokes do_sea() to handle the exception, which calls
>> memory_failure() to handle the faulty page.
>>
>> Scenario 1: Memory Error Interrupt First, then SEA
>> The page is already poisoned by the memory error interrupt path. The
>> subsequent SEA handler sends a SIGBUS to the task, which accesses the
>> poisoned page. This flow is correct.
>>
>> Scenario 2: SEA first, then memory error interrupt (problematic scenario)
>> If a user task directly accesses corrupted memory through a PFNMAP-style
>> mapping (e.g., devmem), the page may still be in the free-buddy state when
>> SEA is handled. In this case, memory_failure() will poison the page without
>> invoking kill_accessing_process(), and then takes the free-buddy recovery
>> path.
>>
>> After the CPU returns to the task context, the task re-enters the SEA
>> handler due to the same access. However, ghes_estatus_cached() suppresses
>> all subsequent entries during the 10-second window, preventing
>> ghes_do_proc() from being called. This suppression blocks the
>> MF_ACTION_REQUIRED-based SIGBUS delivery, causing the kernel to fail to
>> kill the task immediately. Consequently, the process keeps re-entering
>> the SEA handler, leading to an SEA storm. Later, the memory error
>> interrupt path also cannot kill the task, leaving the system stuck in
>> this repeated loop.
>>
>> The following error logs are explained using the devmem process:
>> NOTICE: SEA Handle
>> [Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 9
>> [Hardware Error]: event severity: recoverable
>> [Hardware Error]: section_type: ARM processor error
>> [Hardware Error]: physical fault address: 0x0000001000093c00
>> [T54990] Memory failure: 0x1000093: recovery action for free buddy page: Recovered
>> [ T9955] EDAC MC0: 1 UE Multi-bit ECC on unknown memory
>> (page:0x1000093 offset:0xc00 grain:1 - APEI location: ...)
>> NOTICE: SEA Handle
>> NOTICE: SEA Handle
>> ...
>> ... ---> SEA storm
>> ...
>> NOTICE: SEA Handle
>> [ T9955] Memory failure: 0x1000093: already hardware poisoned
>> ghes_print_estatus: 1 callbacks suppressed
>> [Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 9
>> [Hardware Error]: event severity: recoverable
>> [Hardware Error]: section_type: ARM processor error
>> [Hardware Error]: physical fault address: 0x0000001000093c00
>> [T54990] Memory failure: 0x1000093: already hardware poisoned
>> [T54990] 0x1000093: Sending SIGBUS to devmem:54990 due to hardware memory corruption
>>
>> To resolve this, return an error when encountering the same SEA again.
>> The subsequent SEA handler invocation uses arm64_notify_die() to send a
>> SIGBUS signal to the task, which terminates the process and prevents it
>> from re-entering the handler loop.
>>
>> Signed-off-by: Junhao He <hejunhao3@xxxxxxxxxxxxxx>
>> ---
>> drivers/acpi/apei/ghes.c | 10 +++++++++-
>> 1 file changed, 9 insertions(+), 1 deletion(-)
>>
>> Changes in V2:
>> 1. update the commit message per suggestion from Xueshuai
>> 2. Add a check to only return failure on the ghes_notify_sea() path,
>> avoiding impact on other NMI-type GHES handlers.
>> Link to V1 - https://lore.kernel.org/all/20251030071321.2763224-1-hejunhao3@xxxxxxxxxxxxxx/
>>
>> diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c
>> index 3236a3ce79d6..787664740150 100644
>> --- a/drivers/acpi/apei/ghes.c
>> +++ b/drivers/acpi/apei/ghes.c
>> @@ -1383,8 +1383,16 @@ static int ghes_in_nmi_queue_one_entry(struct ghes *ghes,
> ghes_in_nmi_queue_one_entry() is called from two places:
>
> __ghes_sdei_callback()
>
> I think, but I'm not sure, that you are addressing call from here. By
> returning -ECANCELLED this code skips the irq_work_queue() call and
> returns -ENOENT (instead of zero).
>
> ghes_in_nmi_spool_from_list()
>
> Effect here is more complex. Depends on whether there are multiple
> ghes entries on the rcu_list. If there was just one, and you return
> -ECANCELLED, then irq_work_queue() is skipped. But if there are multiple
> entries on the list and at least one of them gets a zero return from
> ghes_in_nmi_queue_one_entry() then irq_work_queue() is called.
Hi Luck Tony,
Thanks for the comments.
This fix is only for the call to ghes_in_nmi_spool_from_list() from ghes_notify_sea() on
ARM platforms (ACPI_HEST_NOTIFY_SEA); other paths are unchanged.
If multiple SEA errors occur simultaneously, multiple entries in ghes_sea are active
and one of them returns 0, then -ECANCELED is discarded, but the unhandled
SEA will simply trigger again and re-enter the SEA handler to be processed.
The worst case is when multiple valid SEA entries persist within the 10-second
ghes_estatus_cached() window — the fact is that a great many SEA errors are
produced, which actually constitutes a real SEA storm, not one triggered by a single
SEA, and in this situation not delivering SIGBUS in time to kill the process is acceptable.
Best regards,
Junhao.
>
>
> So which of these callers is the one that needs this fix? And is the
> other called OK with getting -ECANCELLED error?
>
>> ghes_clear_estatus(ghes, &tmp_header, buf_paddr, fixmap_idx);
>>
>> /* This error has been reported before, don't process it again. */
>> - if (ghes_estatus_cached(estatus))
>> + if (ghes_estatus_cached(estatus)) {
>> + /*
>> + * Return failure on duplicate SEA entries so that the
>> + * subsequent SEA handler invocation sends a SIGBUS signal to
>> + * the task to prevent it from re-entering the handler loop.
>> + */
>> + if (is_hest_sync_notify(ghes))
>> + rc = -ECANCELED;
>> goto no_work;
>> + }
>>
>> llist_add(&estatus_node->llnode, &ghes_estatus_llist);
>>
>> --
>> 2.33.0
>>
> -Tony
>
> .
>