PROBLEM: fuse: NULL pointer dereference in fuse_dev_do_write() racing device install

From: 马云龙

Date: Fri Oct 09 2026 - 04:24:22 EST


Hi Miklos,

fuse_dev_install_with_pq() publishes fud->chan before it stores
fud->pq.processing. The write/splice_write and read paths treat a
non-NULL fud->chan as "installed" and then walk fpq->processing[]
without taking fch->lock, so they can see the new chan while processing
is still NULL and oops in fuse_request_find().

Introduced by:

48649c0603bd ("fuse: alloc pqueue before installing fch in fuse_dev")

The ordering is unchanged in current mainline.

Details
-------

A fd from open("/dev/fuse") is allocated with fuse_dev_alloc_no_pq(), so
pq.processing == NULL while pq.connected == 1. In
fuse_dev_install_with_pq() (under fch->lock):

old_fch = cmpxchg(&fud->chan, NULL, fch); /* chan published */
...
fud->pq.processing = pq; /* stored afterwards */

fuse_dev_write() only does smp_load_acquire(&fud->chan) via
__fuse_get_dev() and never takes fch->lock:

CPU A: FUSE_DEV_IOC_CLONE CPU B: write(fd2)
------------------------- -----------------
cmpxchg(&fud->chan, NULL, fch)
__fuse_get_dev() sees chan
fuse_dev_do_write()
spin_lock(&fpq->lock)
fuse_request_find()
walks fpq->processing[hash]
-> NULL deref
fud->pq.processing = pq

fuse_dev_release() already takes fch->lock to wait for
fuse_dev_install_with_pq() to finish; the read/write paths have no
equivalent. fuse_dev_do_read() moving a request onto
fpq->processing[hash] has the same exposure.

This looks different from the CUSE double pqueue allocation fix
("fuse: avoid double pqueue allocation in fuse_dev_alloc_install"),
which does not change the chan vs. processing ordering.

Reproducer
----------

Minimal in-guest FUSE server, no libfuse:

1. fd1 = open("/dev/fuse"); mount("fuse", ..., "fd=<fd1>,...").
No sync init, so mount returns without answering FUSE_INIT.
2. Writer threads on other CPUs loop on
write(fd2, fuse_out_header{len=16, error=0, unique=1}).
3. Main thread loops: open a fresh fd2, publish it to the writers,
ioctl(fd2, FUSE_DEV_IOC_CLONE, &fd1), wait for in-flight writes,
close(fd2).

The window is a few instructions inside fch->lock. I hit it on v7.2.9
(run as root) with udelay(200) inserted between the cmpxchg() and the
pq.processing store. That only widens the existing window; locking and
the order of the two stores are unchanged. I have not hit it on an
unmodified kernel. Reproducer source available on request.

Oops excerpt (v7.2.9 + the udelay, 4 vCPUs, KCSAN on):

BUG: kernel NULL pointer dereference, address: 0000000000000000
RIP: 0010:fuse_dev_do_write+0x202/0x7c0
Call Trace:
fuse_dev_write+0x109
vfs_write
ksys_write
__x64_sys_write

faddr2line puts +0x202 at the list_for_each_entry() in
fuse_request_find() (inlined), and +0x18e at the preceding
spin_lock(&fpq->lock). The oopsing task exits with preempt_count 1,
still holding fpq->lock, and another writer then soft-locks in
queued_spin_lock_slowpath from fuse_dev_do_write+0x18e.

Impact: kernel oops plus a stuck fd (DoS). The trigger needs a mounted
FUSE connection plus FUSE_DEV_IOC_CLONE or a racing write; I expect it
is reachable by users who can mount FUSE (fusermount / userns), but I
only tested as root.

Fix direction
-------------

I don't have a tested patch. The obvious direction is to make
pq.processing visible before chan is published. One caveat: two
installs on the same fud can come from different connections (e.g. two
FUSE_DEV_IOC_CLONE calls on the same fresh fd with different source
fds). They hold different fch->lock, so only the cmpxchg() orders them
today. Simply storing processing first and rolling back when the
cmpxchg() fails lets the loser clear and free the queue the winner
ended up with. Serializing installs per fud, or keeping pq.connected at
0 until processing is set (under fpq->lock) so the existing connected
checks reject the half-installed device, might be simpler. I'm happy to
test whatever you prefer.

Thanks,
Yunlong Ma