[PATCH v20 5/9] cxl: Update CXL Endpoint AER handler
From: Terry Bowman
Date: Wed Sep 02 2026 - 10:18:48 EST
Rename cxl_error_detected() to cxl_pci_error_detected() and rename the
struct pci_error_handlers instance from cxl_error_handlers to
cxl_pci_error_handlers for consistency with the cxl_pci_ prefix used by
the renamed handler and the rest of the PCI-facing entry points.
Document the unconditional CXL RAS read policy: on a dead link, readl()
returns 0xFFFFFFFF which is interpreted as UCE bits set and triggers a
panic. If RAS registers are not mapped the read is skipped and the
frozen/perm_failure switch cases defer to AER recovery for devices
without active CXL.mem traffic.
Signed-off-by: Terry Bowman <terry.bowman@xxxxxxx>
Reviewed-by: Dave Jiang <dave.jiang@xxxxxxxxx>
Reviewed-by: Jonathan Cameron <jonathan.cameron@xxxxxxxxxxxxxxxx>
Reviewed-by: Alison Schofield <alison.schofield@xxxxxxxxx>
---
Changes in v19->v20:
- Expand the comment block in cxl_pci_error_detected().
- Clarify commit message's first paragraph. cxl_error_handlers is a variable
- Clarify AER handling in cxl_pci_error_detected() block comment
Changes in v18->v19:
- Add review-by for DaveJ and Jonathan
- Remove reindent introduced in v18 at cxl_error_handlers definition.
Changes in v17->v18:
- Fix cxl_pci_error_detected() to use find_cxl_port_by_uport() and port->uport_dev
- Read CXL RAS unconditionally; panic on UCE regardless of channel state
- Document unconditional read policy and 0xFFFFFFFF behavior in comment
- Drop guard removal paragraph from commit message (not in this diff)
- Drop Reviewed-by tags pending re-review after message change
Changes in v16->v17:
- Rename pci_error_handlers struct instance to cxl_pci_error_handlers for
cxl_pci_ naming consistency.
- Restore scoped_guard(device) and dev->driver check around AER read.
- NULL-check find_cxl_port_by_dev() before deref of port->uport_dev.
- Updated commit message. (Terry)
- Add scope cleanup for port variable in cxl_pci_error_detected() (Terry)
- Drop cxl_uncor_aer_present(), rely on AER state
Changes in v15->v16:
- Update commit message (DaveJ)
- s/cxl_handle_aer()/cxl_uncor_aer_present()/g (Jonathan)
- cxl_uncor_aer_present(): Leave original result calculation based on
if a UCE is present and the provided state (Terry)
- Add call to pci_print_aer(). AER fails to log because is upstream
link (Terry)
Changes in v14->v15:
- Update commit message and title. Added Bjorn's ack.
- Move CE and UCE handling logic here
Changes in v13->v14:
- Add Dave Jiang's review-by
- Update commit message & headline (Bjorn)
- Refactor cxl_port_error_detected()/cxl_port_cor_error_detected() to
one line (Jonathan)
- Remove cxl_walk_port() (Dan)
- Remove cxl_pci_drv_bound(). Check for 'is_cxl' parent port is
sufficient (Dan)
- Remove device_lock_if()
- Combined CE and UCE here (Terry)
Changes in v12->v13:
- Move get_pci_cxl_host_dev() and cxl_handle_proto_error() to Dequeue
patch (Terry)
- Remove EP case in cxl_get_ras_base(), not used. (Terry)
- Remove check for dport->dport_dev (Dave)
- Remove whitespace (Terry)
Changes in v11->v12:
- Add call to cxl_pci_drv_bound() in cxl_handle_proto_error() and
pci_to_cxl_dev()
- Change cxl_error_detected() -> cxl_cor_error_detected()
- Remove NULL variable assignments
- Replace bus_find_device() with find_cxl_port_by_uport() for upstream
port searches.
Changes in v10->v11:
- None
---
drivers/cxl/core/ras.c | 25 ++++++++++++++++---------
drivers/cxl/cxlpci.h | 8 ++++----
drivers/cxl/pci.c | 6 +++---
3 files changed, 23 insertions(+), 16 deletions(-)
diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
index e99dcd028b738..bf479fac08565 100644
--- a/drivers/cxl/core/ras.c
+++ b/drivers/cxl/core/ras.c
@@ -61,7 +61,7 @@ cxl_cper_trace_uncorr_prot_err(struct cxl_memdev *cxlmd,
/*
* ras_cap.header_log[] holds CXL_HEADERLOG_SIZE_U32 (16) hardware
- * dwords. Copy them into the front of a zero-filled
+ * dwords. Copy them into the front of a zero-filled
* CXL_HEADERLOG_TRACE_SIZE_U32 (128) u32 staging buffer so the trace
* event memcpy sees a full 512-byte source and the userspace ABI
* (rasdaemon) is preserved.
@@ -320,8 +320,8 @@ bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem
return true;
}
-pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
- pci_channel_state_t state)
+pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev,
+ pci_channel_state_t state)
{
struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_uport(&pdev->dev);
bool ue = false;
@@ -341,11 +341,18 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
}
/*
- * The CXL RAS read is unconditional regardless of channel
- * state. Any uncorrectable error bit set in the CXL RAS
- * status register triggers a panic below because CXL.mem
- * cache coherency is already lost; continuing risks silent
- * data corruption.
+ * The CXL RAS uncorrectable status is the only signal here
+ * that the error is a CXL internal (protocol) error. A set
+ * UCE bit confirms it and triggers the panic below. On a dead
+ * link readl() returns 0xFFFFFFFF, which sets all UCE bits and
+ * triggers the panic intentionally.
+ *
+ * If RAS is not mapped the read is skipped. Unlike
+ * cxl_do_recovery(), which is reached only after
+ * is_aer_internal_error() has already confirmed a CXL internal
+ * UCE, this path has no such confirmation, so an unmapped RAS
+ * block cannot attribute the error to CXL and must not panic.
+ * The switch cases below then handle AER recovery.
*/
ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL));
}
@@ -373,7 +380,7 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
}
return PCI_ERS_RESULT_NEED_RESET;
}
-EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
+EXPORT_SYMBOL_NS_GPL(cxl_pci_error_detected, "CXL");
static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port,
struct cxl_dport *dport, int severity)
diff --git a/drivers/cxl/cxlpci.h b/drivers/cxl/cxlpci.h
index 606fadd2476f3..7421132899fc4 100644
--- a/drivers/cxl/cxlpci.h
+++ b/drivers/cxl/cxlpci.h
@@ -79,13 +79,13 @@ struct cxl_dev_state;
void read_cdat_data(struct cxl_port *port);
#ifdef CONFIG_CXL_RAS
-pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
- pci_channel_state_t state);
+pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev,
+ pci_channel_state_t state);
void devm_cxl_dport_rch_ras_setup(struct cxl_dport *dport);
void devm_cxl_port_ras_setup(struct cxl_port *port);
#else
-static inline pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
- pci_channel_state_t state)
+static inline pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev,
+ pci_channel_state_t state)
{
return PCI_ERS_RESULT_NONE;
}
diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
index fe11be7fd9aab..1e7be77ded634 100644
--- a/drivers/cxl/pci.c
+++ b/drivers/cxl/pci.c
@@ -998,8 +998,8 @@ static void cxl_reset_done(struct pci_dev *pdev)
}
}
-static const struct pci_error_handlers cxl_error_handlers = {
- .error_detected = cxl_error_detected,
+static const struct pci_error_handlers cxl_pci_error_handlers = {
+ .error_detected = cxl_pci_error_detected,
.slot_reset = cxl_slot_reset,
.resume = cxl_error_resume,
.reset_done = cxl_reset_done,
@@ -1009,7 +1009,7 @@ static struct pci_driver cxl_pci_driver = {
.name = KBUILD_MODNAME,
.id_table = cxl_mem_pci_tbl,
.probe = cxl_pci_probe,
- .err_handler = &cxl_error_handlers,
+ .err_handler = &cxl_pci_error_handlers,
.dev_groups = cxl_rcd_groups,
.driver = {
.probe_type = PROBE_PREFER_ASYNCHRONOUS,
--
2.34.1