Re: [PATCH bpf-next v4 0/2] bpf: htab: Reduce memory use of hash maps
From: T.J. Mercier
Date: Thu Sep 17 2026 - 20:31:56 EST
On Thu, Sep 17, 2026 at 2:21 PM Andrii Nakryiko
<andrii.nakryiko@xxxxxxxxx> wrote:
>
> On Thu, Sep 17, 2026 at 9:31 AM T.J. Mercier <tjmercier@xxxxxxxxxx> wrote:
> >
> > On Thu, Sep 17, 2026 at 9:10 AM Mykyta Yatsenko
> > <mykyta.yatsenko5@xxxxxxxxx> wrote:
> > >
> > > On 8/13/26 12:19 AM, T.J. Mercier wrote:
> > > > Memory is expensive and scarce these days. This series reduces the
> > > > memory use of BPF hash maps by eliminating the per-element overheads
> > > > below. This saves up to 50% of per-element memory use for standard and
> > > > PCPU hash maps. The memory use of LRU hash maps is unaffected.
> > > >
> > > > Map Type & Configuration | Old size | New size | Savings
> > > > ------------------------------------|----------|----------|--------
> > > > Standard (key <= 8 B, val <= 8 B) | 64 B | 32 B | 50.0%
> > > > Per-CPU (prealloc) (key <= 8 B) | 64 B | 32 B | 50.0%
> > > > Per-CPU (non-prealloc) (key <= 8 B) | 64 B | 40 B | 37.5%
> > > > LRU (Any key/value size) | - | - | 00.0%
> > > >
> > >
> > > T.J. are you still interested landing this? Maybe respin the series?
> > > Alexei was away back when you sent this.
> >
> > Hi Mykyta, thanks for following up. Yes, I'd still like to land it. I
> > took a break during the merge window which coincided with a vacation,
> > where I broke my collarbone on a mountain bike jump gone wrong. Then
> > surgery and recovery, and I'm still catching up on everything from
>
> oh, wow, hope you'll heal fast and well!
Thanks! The surgery already made a huge difference, so now I'm waiting
for bone to grow and it'll be a few months before I'm back on a bike.
> > while I was out. I plan to rebase this and send it out again before I
> > head out for pre-LPC travel next week.
>
> have you considered splitting lru and non-lru flavors of hashtable
> before doing this optimization? I'm wondering if it will allow to
> clean up some parts of it, while also making map struct itself smaller
> for non-lru map (there is that LRU-specific piece in the union which
> artificially blows up the size of any hash map).
>
> I am a bit concerned about that key offset, even if the microbenchmark
> doesn't show much difference. What if we put hash itself before
> per-element header, so that key/value are always at the same offset.
> For cases where key size > 8 we'll just access hash at
> addrof(htab_elem) - 8, while smaller key sizes will just directly
> compare keys.
This is an interesting idea and I think it will work. Let me try this
first, and then I will take a look at splitting out LRU hashtables
aftewards.
> Just some high level thoughts, haven't really coded any of that, so
> hard to tell upfront if it's worth doing.
>
> >
> >
> > > > 1) Unused LRU / PCPU fields in standard and PCPU hash maps (patch 1)
> > > > struct htab_elem is used for all hash map types, and includes fields
> > > > that are not always used (bpf_lru_node, ptr_to_pptr). For standard
> > > > (non-LRU, non-PCPU) hash maps the 24 bytes for the bpf_lru_node (union)
> > > > are entirely overhead and can be eliminated. Non-preallocated PCPU maps
> > > > only need the 8 byte ptr_to_pptr which is currently unioned with the
> > > > unneeded 24 byte bpf_lru_node, so 16 bytes of overhead can be
> > > > eliminated. Preallocated PCPU maps don't need ptr_to_pptr, so 24 bytes
> > > > of overhead can be saved.
> > > >
> > > > 2) Hash caching for small keys (patch 2)
> > > > For hash maps with small key sizes (<= word size), comparing keys only
> > > > requires a single instruction. Currently the 4 byte hash value (8 byte
> > > > aligned and padded) is used for this, but offers no performance
> > > > advantage in this case and can be eliminated.
> > > >
> > > > The implementation splits htab_elem into dedicated structures for the
> > > > different map types (htab_elem_lru, htab_elem_pcpu, htab_elem) which
> > > > share a common initial sequence (struct htab_node), but contain
> > > > additional map-type specific fields where necessary. This means the
> > > > placement of the key for each element varies with the map type, and
> > > > key_offset is added to bpf_htab for this purpose.
> > > >
> > > > While using key_offset and conditional hash checks adds new pointer
> > > > dereferences and branching during element traversal,
> > > > run_bench_htab_mem.sh shows no significant performance regression across
> > > > 10 runs on my 3995WX.
> > > >
> > > > Benchmark (all in kops/sec) | Avg. Before | StDev | Avg. After | StDev
> > > > -----------------------------|-------------|-------|------------|------
> > > > prealloc overwrite | 115.11 | 4.10 | 115.45 | 5.24
> > > > prealloc batch_add_batch_del | 127.14 | 4.32 | 127.06 | 2.32
> > > > prealloc add_del_on_diff_cpu | 23.22 | 0.93 | 22.91 | 1.60
> > > > normal overwrite | 78.52 | 3.05 | 80.40 | 3.25
> > > > normal batch_add_batch_del | 45.37 | 0.69 | 47.71 | 0.66
> > > > normal add_del_on_diff_cpu | 12.02 | 0.73 | 12.48 | 0.70
> > > >
> > > > ---
> > > > Changes in v4:
> > > > Removed inline from new functions per BPF CI (netdev/source_inline).
> > > >
> > > > From Mykyta Yatsenko:
> > > > Factor out duplicate lookup_elem code into __lookup_elem_raw.
> > > > Use offsetof instead of sizeof for key_offset assignments in
> > > > htab_map_alloc (patch 1).
> > > > Eliminate branching and htab_elem casting in htab_elem_hash /
> > > > htab_elem_set_hash.
> > > >
> > > > Changes in v3:
> > > > From Sashiko on torn reads/writes:
> > > > Use a local unsigned long and READ_ONCE / WRITE_ONCE instead of memcmp /
> > > > memcpy for atomic key comparisons for hashless elements.
> > > >
> > > > Changes in v2:
> > > > Make maximum key_size for !has_hash depend on word size for atomicity
> > > > on 32-bit.
> > > >
> > > > From Mykyta Yatsenko:
> > > > Put the htab_elem* common initial sequence in its own struct (htab_node)
> > > > and reuse it across all element types that share it. Eliminate
> > > > associated BUILD_BUG_ON additions.
> > > > Replace both the hash and key fields with data[].
> > > > Store has_hash in struct bpf_htab, and avoid per-element reads of it.# Please edit the description for the branch
> > > >
> > > > T.J. Mercier (2):
> > > > bpf: htab: Split htab_elem_lru and htab_elem_pcpu off of htab_elem
> > > > bpf: htab: Reduce elem_size by 8 bytes for small key sizes
> > > >
> > > > kernel/bpf/hashtab.c | 418 ++++++++++++------
> > > > kernel/bpf/map_in_map.c | 13 +
> > > > kernel/bpf/map_in_map.h | 2 +
> > > > .../selftests/bpf/progs/map_ptr_kern.c | 2 +-
> > > > 4 files changed, 289 insertions(+), 146 deletions(-)
> > > >
> > > >
> > > > base-commit: cfce77b63375dac81d53f2f85593c548415206b7
> > >