Re: mlx5e:zero-prefix corruption during TCP DMA-BUF RX - help requested
From: jiabin deng
Date: Thu Oct 08 2026 - 02:49:51 EST
Hi Dragos,
Thank you for following up and suggesting the CPU-buffer comparison. We have
now completed both ordinary CPU TCP and CPU-backed DMA-BUF receive controls,
with no payload corruption in either 400-iteration run.
Please find attached a standalone GPU-memory reproducer with a README covering
the required environment, build instructions, two-host execution, and result
collection. It includes a Python TCP sender, a C++ receiver using the CUDA
Driver API and Linux device-memory TCP interfaces, and scripts for SHAMPO
header tracing and offline byte comparisons.
As reported in my previous reply, the corruption was still reproducible after we
upgraded the receiver from Ubuntu kernel `7.0.0-30-generic` to Linux
`7.2.6-070206-generic`, using the in-tree mlx5_core driver. The v7.2.6 source
contains the small-frame guard associated with commit
[e2466392a0b8496000e12181cb1ee1535eb0da25](https://github.com/torvalds/linux/commit/e2466392a0b8496000e12181cb1ee1535eb0da25).
The attached program retains the receive, synchronization, and data-comparison
sequence used in those experiments.
The reproduction sequence is:
1. Enable `tcp-data-split` and `rx-gro-hw`, allocate and export 4 GiB of GPU
memory as a DMA-BUF, bind it to RX queues 8–15, and steer the test flow to
queue 8. The allocation and binding stay live throughout the sweep.
2. Before each iteration, prefill GPU memory with a nonzero value and complete
`cuCtxSynchronize()`. A receiver-to-sender ready byte ensures the sender
does not transmit the payload before this step completes.
3. Send deterministic data whose byte values are all in the range 1–255.
Receive the complete message and EOF with `MSG_SOCK_DEVMEM`, recording
the fragment offsets, lengths, DMA-BUF IDs, and tokens.
4. Retain every token while saving a full-message snapshot with
`cuMemcpyDtoH()`, calling
`cuFlushGPUDirectRDMAWrites(CURRENT_CTX, TO_ALL_DEVICES)`, and saving a
second snapshot.
5. Compare the snapshots with each other and the first snapshot with the
expected data. Return the tokens through `SO_DEVMEM_DONTNEED` only after
the reads, snapshot writes, and comparisons have completed.
The default test uses eight message sizes from 32 KiB to 4 MiB, with 50
sequential connections per size. The link MTU is 1500, and the sender checks
that the negotiated TCP MSS is 1448 bytes. The README explains how to select
the receiving NIC interface, IP addresses, and GPU PCI bus ID for your test
environment, and lists the expected readiness and completion markers. Header
tracing uses bpftrace and the running kernel's BTF type information.
For the CPU controls, we used the same sender, deterministic nonzero payload
pattern, eight message sizes, 50 iterations per size, and negotiated MSS of
1448 bytes. In both controls, `tcp-data-split` and `rx-gro-hw` remained enabled,
and the test flow was steered to RX queue 8:
- Ordinary CPU TCP: receive with `recvmsg(..., 0)` into an application buffer
of up to 4 MiB, without binding a DMA-BUF to the RX queues.
- CPU-backed DMA-BUF: export a 4 GiB system-RAM pool through `udmabuf`, bind it
to RX queues 8–15, and receive with `MSG_SOCK_DEVMEM`. The pool was prefilled
with a nonzero value before each iteration. We retained the tokens through
both full-message snapshots and byte comparisons, then returned them with
`SO_DEVMEM_DONTNEED`. CPU accesses were bracketed by
`DMA_BUF_IOCTL_SYNC` START/END with RW flags.
The two full-message snapshots were identical in every iteration of both
controls, and every byte matched the expected payload.
Given that neither CPU control reproduced the corruption seen with GPU-backed
DMA-BUF reception, what additional checks would you recommend for the GPU case?
Please also let us know if you spot any issues with the reproducer's receive
logic, token management, or synchronization.
Thanks,
Jiabin Deng
On Tue, Sep 29, 2026 at 5:49 PM Dragos Tatulea <dtatulea@xxxxxxxxxx> wrote:
>
>
>
> On 18.09.26 08:06, jiabin deng wrote:
> > Hi Dragos,
> >
> > Thank you for pointing us to this fix.
> >
> > Our original test used Ubuntu kernel 7.0.0-30-generic, based on
> > v7.0.12, and did not include that change. We have now repeated the
> > test using Ubuntu's mainline build of Linux 7.2.6
> > (7.2.6-070206-generic), with its in-tree mlx5_core driver.
> >
> > I checked the v7.2.6 source: it contains the small-frame guard from
> > commit e2466392a0b8496000e12181cb1ee1535eb0da25 In
> > mlx5e_handle_rx_cqe_mpwrq_shampo(), frames with
> > cqe_bcnt <= ETH_ZLEN + 2 * VLAN_HLEN set match = false and
> > flush = true [2].
> >
> > The zero-filled corruption ending at the next 64-byte boundary in
> > GPU memory still reproduces with this kernel.
> >
> > We kept the same sender and receiver data-checking logic, with
> > tcp-data-split on and rx-gro-hw on. One additional environment
> > change is that the NVIDIA GPU driver was updated from 595.71.05 to
> > 595.91.07. The NIC firmware remains 32.43.2400 (MT_0000001117).
> >
> > We repeated the same eight message sizes, from 32 KiB to 4 MiB,
> > with 50 iterations per size:
> >
> > - All 400 iterations completed.
> > - 105 iterations passed; 295 contained corrupted data.
> > - 21,599 corrupted ranges were recorded, all zero-filled and
> > ending at the next 64-byte boundary in GPU memory.
> > - For each size from 256 KiB through 4 MiB, all 50 iterations
> > contained corruption.
> > - The full-message snapshots taken before and after
> > cuFlushGPUDirectRDMAWrites() were identical in all iterations.
> >
> > For example, in the first 32 KiB iteration, an 8-byte corrupted
> > range starts at offset 2147495416 bytes from the beginning of the
> > GPU buffer. This offset is 56 bytes into a 64-byte-aligned block,
> > so the eight zero bytes end exactly at the next 64-byte boundary.
> >
> > These results come from the reproducer's runtime byte comparisons.
> >
> > Could you suggest the next targeted checks or additional debug
> > information that would help narrow this down?
> >
> Did you also try CPU buffers? From the NIC perspective it shouldn't
> be any different.
>
> Could you share a program + reproduction steps?
>
> Thanks,
> Dragos
Attachment:
mlx5_cuda_dmabuf_reproducer.zip
Description: Zip archive