Re: [PATCH net 1/2] lib/dim: fix 32-bit overflow in dim_calc_stats() rates
From: shashank Jain
Date: Wed Sep 30 2026 - 05:43:07 EST
> How does this 32-bit overflow differ from the overflow that can also occur on
> 64-bit systems in `nbytes * USEC_PER_MSEC`?
On 64-bit the multiplication itself cannot overflow. nbytes (like npkts
and ncomps) is a u32, and USEC_PER_MSEC is 1000L, so the product is
computed in a 64-bit long and is at most (2^32 - 1) * 1000, about
4.3 * 10^12 (42 bits). DIV_ROUND_UP() adds at most delta_us - 1 to that,
which also fits. So on 64-bit the product is always exact.
On 32-bit the same product is computed in a 32-bit long and wraps once
nbytes exceeds 4294967, i.e. about 4.3 MB in one DIM window. That is
reached at normal line rates, which is what the patch fixes. With the
u64 cast and DIV_ROUND_UP_ULL() the 32-bit results are the same as on
64-bit.
There are two other limits, but they apply to 32-bit and 64-bit alike
and the patch does not change them:
- The counters in struct dim_sample are u32 and BIT_GAP() takes the
difference modulo 2^32, so a window that carries 4 GiB or more is
already undercounted before the multiplication. With 64 events per
window that needs a very long window at very high rates (for example
about 86 ms at 400 Gbit/s).
- The rates are stored in int fields of struct dim_stats. bpms is bytes
per millisecond, so it only exceeds INT_MAX above 2^31 bytes/ms,
about 17 Tbit/s.
I can add a sentence to the changelog saying that the 64-bit product
cannot overflow, if you think that helps.
Thanks,
Shashank
On Wed, Sep 30, 2026 at 2:19 PM Leon Romanovsky <leon@xxxxxxxxxx> wrote:
>
> On Sun, Sep 27, 2026 at 10:47:42AM +0530, Shashank Mohan Jain wrote:
> > dim_calc_stats() computes the per-millisecond rates as
> >
> > DIV_ROUND_UP(nbytes * USEC_PER_MSEC, delta_us)
> >
> > where nbytes is a u32 and USEC_PER_MSEC is 1000L. On 64-bit the product
> > is done in 64-bit long arithmetic, but on 32-bit architectures long is
> > 32 bits wide and the product wraps as soon as a measurement window
> > carries more than 4294967 bytes (about 4.3 MB). The same applies to the
> > packet and completion counts, although those need more than 4.29
> > million packets or completions per window.
> >
> > A DIM window spans DIM_NEVENTS (64) events. Drivers count events per
> > interrupt or per NAPI poll, so under sustained load a window can easily
> > carry more than 4.3 MB: 64 full NAPI polls of 64 MTU-sized frames are
> > already 6.2 MB, and drivers such as mtk_eth_soc count one event per
> > interrupt while NAPI keeps polling with the interrupt masked. On 32-bit
> > users of the library (for example mtk_eth_soc on MT7621, bcmgenet and
> > bcmsysport on 32-bit ARM, or virtio_net in a 32-bit guest) bpms then
> > becomes the product modulo 2^32 divided by the window length, and
> > net_dim_stats_compare() makes its BETTER/WORSE decisions on a value
> > that has little to do with the real throughput.
> >
> > For example, a 1 Gbit/s link at line rate that moves 5 MB in a 40 ms
> > window gives bpms = 125000 on 64-bit but 17626 on 32-bit, and 5 million
> > packets in 2 s gives ppms = 353 instead of 2500.
> >
> > Widen the products to 64 bits and divide with DIV_ROUND_UP_ULL(). The
> > results are unchanged on 64-bit.
>
> How does this 32-bit overflow differ from the overflow that can also occur on
> 64-bit systems in `nbytes * USEC_PER_MSEC`?
>
> Thanks