Re: [PATCH v11 4/7] wifi: rtw88: sdio: zero the padding added to a TX transfer
From: Luka Gejak
Date: Fri Sep 11 2026 - 03:42:38 EST
On Fri Sep 11, 2026 at 2:53 AM CEST, Ping-Ke Shih wrote:
> Ping-Ke Shih <pkshih@xxxxxxxxxxx> wrote:
>> Luka Gejak <luka.gejak@xxxxxxxxx> wrote:
>> > mac80211 reserves IEEE80211_ENCRYPT_TAILROOM, 18 bytes, and offers
>> > no way for a driver to ask for more on TX; extra_tx_headroom is headroom
>> > and extra_beacon_tailroom is beacons only.
>>
>> You can modify mac80211 to reserve larger ndev->needed_tailroom to
>> see if it can really resolve the symptom. If so, you can propose
>> to add fields for tailroom like headroom:
>>
>> struct ieee80211_hw:: extra_tx_tailroom
>> struct ieee80211_local:: tx_headroom
>>
>
> Note that the ndev->needed_tailroom isn't guaranteed by comments, so doing
> some experiments by normal use case would be helpful to know if it's worth.
>
> * @needed_tailroom: Extra tailroom the hardware may need, but not in all
> * cases can this be guaranteed. Some cases also use
> * LL_MAX_HEADER instead to allocate the skb
>
> Ping-Ke
I ran it: one line in net/mac80211/iface.c, IEEE80211_ENCRYPT_TAILROOM
+ 512, nothing else. It does not help. Over 60000 frames the average
tailroom went from 106 to 139 bytes against an average pad of 469, and
96% still reallocate.
Your caveat is why, and it is stronger than the comment suggests:
mac80211 never reads ndev->needed_tailroom as the field appears once in
net/mac80211, which is the assignment itself. The only code that acts on
it is skb_ensure_writable_head_tail(), whose one caller is net/dsa/user.c.
And the protocols that do honour it read it when they allocate, which TCP
never does, which is why only the few non-TCP frames moved the average.
So that tested the wrong knob rather than the idea.
Your fields do work. I built them and measured it. ieee80211_skb_resize()
is the only place on the TX path that grows tailroom, and it derives
tail_need from IEEE80211_ENCRYPT_TAILROOM alone, gated on the frame
needing software crypto tailroom, which is false for data frames under
hardware CCMP. The headroom side is already what you describe:
hw.extra_tx_headroom folds into local->tx_headroom and both callers add
it into head_need. Adding hw.extra_tx_tailroom, folding it into a new
local->tx_tailroom and adding that to tail_need outside the crypto gate
delivers the tailroom. One further piece is needed: ieee80211_build_hdr()
skips the resize entirely when it wants no headroom and the skb is not
cloned, so that condition has to widen as well.
With the driver asking for one SDIO block, frames arriving short of
tailroom go from 29056 in 30000 to 1 in 30000, and uplink is about 6%
faster in an interleaved A/B.
One caveat though: this does not remove the reallocation, it moves it.
pskb_expand_head() is called just as often, from ieee80211_skb_resize()
now instead of from __skb_pad() in the driver. The driver side becomes
much cheaper, 445 ms of __skb_pad per 20 s run against 82 ms, and total
time in pskb_expand_head falls by about a third, but whole system CPU
does not drop. The gain looks like moving the work off the SDIO critical
path rather than doing less of it.
I can write that up for Johannes as a separate mac80211 patch, with the
rtw88 side as a follow up. It is independent of this series, but considering
your question was about CPU usage and this change doesn't overall reduce it,
it is probably not worth it.
Best regards,
Luka Gejak