Re: [PATCH net v4] net: mana: fix reset work race with device removal
From: Fan Wu
Date: Sun Sep 27 2026 - 06:51:23 EST
Thanks for the review. All three new findings are legitimate; v5
(following as a separate thread) addresses them:
- [Medium] mana_gf_stats_work_handler(): agreed. When
mana_rdma_probe() outlives the 2s stats period, the gate dropped the
reset request with no latch and no re-arm, so a probe that then
succeeded never recovered the stats work. v5 takes the re-arm
option: the gated branch re-arms the delayed work, retrying once
the probe completes; a failing probe still cancels it in the
mana_remove() unwind. The latch alternative would roll back a
healthy probe on a single transient timeout, so I left it out.
- [Low] system_wq: agreed, and the work also escaped
flush_scheduled_work(), since system_wq is a separate compatibility
queue. v5 returns to schedule_work(), the pre-patch
system_percpu_wq placement. Whether a long queue suits the 10s+
service cycles better can be a follow-up of its own.
- [Low] service_quiesce comment: agreed. It now states that only the
rollback path can have admitted a cycle, because it runs
mana_service_probe_complete() before jumping there.
On the [High]: the PM suspend/resume path is outside this series'
remove/free closure, and applying quiesce there directly would close
admission one-way across suspend/resume and would deadlock a service
cycle that is itself calling mana_gd_suspend(). A reversible wait
plus a resume-side reopen is needed, which will follow as a separate
patch.
pw-bot: cr