Re: [PATCH] accel/rocket: search every core slot when a core is removed

From: Sidong Yang

Date: Sat Sep 05 2026 - 09:28:29 EST


On Fri, Sep 04, 2026 at 02:59:36PM +0200, Igor Paunovic wrote:
> rocket_remove() decrements rdev->num_cores for each core it removes,
> while find_core_for_dev() searches slots 0 to num_cores - 1. Unbinding
> the cores in the order they were bound therefore loses the last one: by
> the time it is removed the search range has already shrunk past its
> slot, so find_core_for_dev() returns -1 and rocket_remove() gives up
> without doing anything.
>
> num_cores never reaches zero, rocket_device_fini() never runs, and the
> file-scoped rdev keeps pointing at a device that is going away. Binding
> the cores again starts from that stale count, because rocket_probe()
> takes rdev->num_cores as the slot to fill. On an RK3588, which describes
> three cores, the second round lands on slots 1, 2 and 3 while rdev->cores
> was allocated with room for three:
>
> rocket fdab0000.npu: Rockchip NPU core 1 version: 1179210309
> rocket fdac0000.npu: Rockchip NPU core 2 version: 1179210309
> rocket fdad0000.npu: Rockchip NPU core 3 version: 1179210309
>
> The write to rdev->cores[3] is past the end of the array.
>
> Nothing in tree reads the core array often enough to notice, so the
> overrun is silent today. It turned up while testing a devfreq series on
> top of this, where a worker walks every core a few times a second, and
> UBSAN caught the first bool it read out of the overrun entry:
>
> UBSAN: invalid-load in drivers/accel/rocket/rocket_devfreq.c:47:10
> load of value 5 is not a valid value for type '_Bool'
> Workqueue: devfreq_wq devfreq_monitor
>
> Record how many slots were allocated and search all of them. Every core
> is then found on removal, num_cores reaches zero, the device is torn down
> and a later bind starts from a clean rdev.
>
> This does not make unbinding a single core out of several work. probe
> still takes num_cores as the slot to fill, so rebinding one core while
> its siblings stay bound would write over a slot that is already in use,
> and rocket_open() still reaches for cores[0] whether or not anything is
> there. Both of those want more thought than a fix should carry.
>
> Found by unbinding and rebinding all three cores on an Orange Pi 5 Plus.
> With this applied, 25 unbind/rebind rounds and 5 module unload/reload
> rounds run clean there: the cores land in slots 0, 1 and 2 every time,
> whichever order they are bound in, and the shared supply goes back to a
> single user after each round.
>
> Fixes: ed98261b4168 ("accel/rocket: Add a new driver for Rockchip's NPU")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Igor Paunovic <royalnet026@xxxxxxxxx>
> Assisted-by: LLM sparse checkpatch
> ---
> drivers/accel/rocket/rocket_device.c | 2 ++
> drivers/accel/rocket/rocket_device.h | 1 +
> drivers/accel/rocket/rocket_drv.c | 2 +-
> 3 files changed, 4 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/accel/rocket/rocket_device.c b/drivers/accel/rocket/rocket_device.c
> index 46e6ee1e72c5f..efd004194c1af 100644
> --- a/drivers/accel/rocket/rocket_device.c
> +++ b/drivers/accel/rocket/rocket_device.c
> @@ -31,6 +31,8 @@ struct rocket_device *rocket_device_init(struct platform_device *pdev,
> if (of_device_is_available(core_node))
> num_cores++;
>
> + rdev->max_cores = num_cores;
> +
> rdev->cores = devm_kcalloc(dev, num_cores, sizeof(*rdev->cores), GFP_KERNEL);
> if (!rdev->cores)
> return ERR_PTR(-ENOMEM);
> diff --git a/drivers/accel/rocket/rocket_device.h b/drivers/accel/rocket/rocket_device.h
> index ce662abc01d3d..c62d567010696 100644
> --- a/drivers/accel/rocket/rocket_device.h
> +++ b/drivers/accel/rocket/rocket_device.h
> @@ -19,6 +19,7 @@ struct rocket_device {
>
> struct rocket_core *cores;
> unsigned int num_cores;
> + unsigned int max_cores;

Hi Igor,

I've tested this patch in Radxa Rock 5 B+ and it works. The test is that unbinding all
cores in sequence 0,1,2 and rebind all. It seems that there is other issue about num_core.
For example, sched_to_core() finds core for sched with num_core and it could make same
error like find_core_for_dev().

Tested-by: Sidong Yang <sidong.yang@xxxxxxxxxx> # RK3588, 3 cores

Thanks,
Sidong

> };
>
> struct rocket_device *rocket_device_init(struct platform_device *pdev,
> diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocket_drv.c
> index 8bbbce594883e..2bcfe4ab3c68f 100644
> --- a/drivers/accel/rocket/rocket_drv.c
> +++ b/drivers/accel/rocket/rocket_drv.c
> @@ -223,7 +223,7 @@ static int find_core_for_dev(struct device *dev)
> {
> struct rocket_device *rdev = dev_get_drvdata(dev);
>
> - for (unsigned int core = 0; core < rdev->num_cores; core++) {
> + for (unsigned int core = 0; core < rdev->max_cores; core++) {
> if (dev == rdev->cores[core].dev)
> return core;
> }
> --
> 2.43.0
>