Re: [PATCH 1/9] EDAC/versalnet: Add NULL check for mci in handle_error()
From: Pandey, Radhey Shyam
Date: Wed Jul 29 2026 - 13:38:11 EST
On 7/29/2026 3:03 AM, Borislav Petkov wrote:
On Tue, Jul 28, 2026 at 11:57:28PM +0530, Pandey, Radhey Shyam wrote:
One clarification on DEV_UNKNOWN: on Versal NET, the driver exposes 8
MC5 controller slots, but a given platform may not have all of them
configured.
Are you basically saying that you can have non-contiguous controller slots
present on a system?
In the Versal NET designs I've examined so far, configuration with fewer
than eight MC5 controllers uses a contiguous set starting at slot 0
(for example, slots 0-3 populated and slots 4-7 absent).
I will let Shubhrajyoti comment if non-contiguous controller slots
are possible on any supported platform?
Agreed for real init failures if initialization of a controller that
should be present fails, probe should fail and unwind all instances
registered so far, same as amd64_edac.
For v2, my proposal is (only if contiguous slot are supported):
-Take the supported controller count from the design/DT rather than
always iterating hardcoded NUM_CONTROLLERS(8). Though it has dependency
on XSA support(hardware description metadata archive exported from
Vivado) and DT binding getting accepted.
-Run init/remove only for that range.
-Treat DEV_UNKNOWN as a probe error with full unwind, not as a silent
skip. Indices beyond the supported count are not probed at all.
-Document this model in the driver.
-The existing WARN_ON_ONCE() check in handle_error() should be
sufficient so the current patch can be dropped.
Shubhrajyoti: please chime in if I've missed anything or if you
have a different view on this approach.
Thanks,
Radhey