Re: [RFC PATCH 3/3] io_uring/rsrc: prefill the node cache when a file table is registered empty
From: Uzair Beg
Date: Fri Sep 25 2026 - 12:42:19 EST
Gabriel Krisman Bertazi <krisman@xxxxxxx> writes:
> I'm unconvinced this is the right approach. This is only relevant for
> initialization overhead: once the ring is in operation, the overhead is
> gone because nodes are recycled.
For the first fill I agree. The part I would push back on is the
recycling. The node cache holds 128 entries, so an application that
removes and reinstalls more than 128 files keeps allocating past that
point. In the remove and refill test from the cover letter (4096 files
on one ring, FILES_UPDATE with -1 followed by SEND_FD) the unpatched
kernel allocated 3968 new nodes per cycle. As far as I can tell the
18.8% there comes from the larger cache keeping those nodes rather
than from the prefill itself, since the prefill runs once at
registration, outside the timed cycles.
> So this will really benefit
> short-lived applications that create large tables, something that I
> suspect is rare outside of artificial benchmarks.
That is fair. The reported workload is a microbenchmark, and I don't
have a real application to point to that fills a large sparse table
once and exits.
> On the other hand,
> people are creating sparse but arbitrarily large tables. Does it make
> sense to pre-allocated up to 192KB in memory for short-lived
> applications that might use only a couple of those nodes?
As a default, I don't think it does. Would it be acceptable to drop the
prefill and only let the node cache capacity follow the table size,
capped? Nothing would be allocated up front beyond the pointer array,
so registering a large table and using a few slots costs almost
nothing, and nodes are only retained up to what the application
actually had installed. That gives up the first fill result but keeps
the churn case, and it needs neither the dedicated slab nor the bulk
refill.
I'll measure that variant and follow up with numbers before posting
anything. If you would rather the node cache stay at a fixed size,
that is useful to know too.