Re: [PATCH v2] wifi: ath11k: release peer accounting on peer delete timeout

From: Baochen Qiang

Date: Tue Sep 29 2026 - 02:20:16 EST




On 9/24/2026 5:14 PM, Michael Pfeifroth wrote:
> On some deployments access points intermittently stop accepting new
> station associations after several hours of uptime with frequent
> roaming/reconnects. The kernel logs
>
> ath11k_pci ....: failed to create peer due to insufficient peer entry resource in firmware
>
> and hostapd reports "Could not add STA to kernel driver". A "wifi
> down/up" (radio restart) on the affected radio restores service.
>
> Despite the message text, this is not a firmware peer-table exhaustion.
> The message is emitted by the driver-side gate in ath11k_peer_create():
>
> if (ar->num_peers > (ar->max_num_peers - 1))
>
> i.e. the driver's own ar->num_peers accounting has leaked and reached
> the ceiling. It is always preceded by a peer-delete that timed out:
>
> ath11k_pci ....: invalid vdev id in peer delete resp ev 1
> ath11k_pci ....: Timeout in receiving peer delete response
> ath11k_pci ....: failed to delete peer vdev_id .. addr .. ret -110
>
> ath11k_peer_delete() only decrements ar->num_peers when
> __ath11k_peer_delete() returns 0. On a delete timeout
> __ath11k_peer_delete() returns -ETIMEDOUT, so the decrement is skipped
> and one num_peers slot is leaked per event. After max_num_peers such
> timeouts ath11k_peer_create() rejects every new station until the radio
> is restarted.
>
> In the observed case the peer-unmap event is received normally (there is
> no "failed wait for peer deleted" log), so ath11k_peer_unmap_event() has
> already removed the peer from ab->peers and freed it; only the num_peers
> counter is left wrong. The delete-response completion is missed because
> the response event is dropped in ath11k_peer_delete_resp_event() when
> ath11k_mac_get_ar_by_vdev_id() cannot resolve the vdev ("invalid vdev id
> in peer delete resp ev"), so ar->peer_delete_done is never signalled and
> the second wait in ath11k_wait_for_peer_delete_done() times out.
>
> Return success from __ath11k_peer_delete() on the timeout path so that
> ath11k_peer_delete() releases the num_peers slot. As a safety net also
> drop the local peer if it is still on the list; that only happens in the
> other timeout case, where ath11k_wait_for_peer_deleted() itself timed out
> and no unmap event removed the peer. The peer has already been removed
> from the rhash earlier in __ath11k_peer_delete(), so only the list
> removal and free remain, mirroring ath11k_peer_unmap_event(). A late
> unmap event would then simply fail to find the peer id and log a harmless
> warning instead of touching freed memory.
>
> The problem was reproduced deterministically with an out-of-tree debug
> patch that adds module parameters to force
> ath11k_wait_for_peer_delete_done() to return -ETIMEDOUT and to cap
> max_num_peers: after max_num_peers such deletes the AP permanently
> rejects new stations, and with this change it keeps accepting them.

though the last paragraph explicitly says the issue is artificial, most of the commit
message still reads like a real world bug ...

I think it would be better to rephrase as something like 'Theoretically, if peer unmap
event is good but peer delete fails ...'

>
> Fixes: 690ace20ff79 ("ath11k: peer delete synchronization with firmware")
> Cc: stable@xxxxxxxxxxxxxxx

I don't think it qualifies for stable backport since it is not a real world issue.

> Signed-off-by: Michael Pfeifroth <micpf@xxxxxxxxxxxx>
> ---
> v2:
> - Correct the root-cause description: the peer-unmap event is received
> normally, so the peer is already freed; only the num_peers counter
> leaks (Baochen Qiang).
> - Explain that the missed delete-response completion is due to the event
> being dropped on an unresolved vdev id, replacing the vague
> "misrouted" wording.
> - Rework the code comment and switch to netdev comment style.
> drivers/net/wireless/ath/ath11k/peer.c | 17 +++++++++++++++--
> 1 file changed, 15 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/net/wireless/ath/ath11k/peer.c b/drivers/net/wireless/ath/ath11k/peer.c
> index b30a906..f93a8da 100644
> --- a/drivers/net/wireless/ath/ath11k/peer.c
> +++ b/drivers/net/wireless/ath/ath11k/peer.c
> @@ -341,8 +341,21 @@ static int __ath11k_peer_delete(struct ath11k *ar, u32 vdev_id, const u8 *addr)
> }
>
> ret = ath11k_wait_for_peer_delete_done(ar, vdev_id, addr);
> - if (ret)
> - return ret;
> + if (ret) {
> + /*
> + * The delete timed out. The peer is normally already freed by
> + * the unmap event; drop it here only if it is still on the
> + * list. Either way return success so that ath11k_peer_delete()
> + * releases the num_peers slot instead of leaking it.
> + */
> + spin_lock_bh(&ab->base_lock);
> + peer = ath11k_peer_find(ab, vdev_id, addr);
> + if (peer) {
> + list_del(&peer->list);
> + kfree(peer);
> + }
> + spin_unlock_bh(&ab->base_lock);

except for the firmware crash case, this is exactly what is done in
ath11k_wait_for_peer_deleted(), so if that function succeeds the peer is definitely
removed and freed already, so this is actually dead code.

as for the crash case, all peers are cleaned up in recovery path so we don't need to do
remove/free as well.

> + }
>
> return 0;
> }