Re: [PATCH v2] fs/locks: filter OFD locks from /proc/locks for foreign pid namespaces

From: Tao Cui

Date: Wed Sep 02 2026 - 22:14:07 EST


Hi, Jeff

在 2026/9/2 21:03, Jeff Layton 写道:
> On Wed, 2026-09-02 at 20:51 +0800, Tao Cui wrote:
>> From: Tao Cui <cuitao@xxxxxxxxxx>
>>
>> Since Linux 4.9, /proc/locks only shows locks whose owning process is
>> visible in the reader's pid namespace. The check cannot see the
>> owner of an OFD lock, though: locks_translate_pid() always returns
>> -1 for FL_OFDLCK, so an OFD lock held outside of that namespace
>> bypasses the filter and is shown with its device/inode numbers and
>> byte range, e.g.
>>
>> 4: OFDLCK ADVISORY WRITE -1 fd:00:70536961 0 EOF
>>
>> This was observed on a Kubernetes node: a pod reading /proc/locks
>> listed the OFD write lock of an unrelated host process, while the
>> POSIX locks of the very same process were correctly hidden. The
>> device/inode pair identifies which host file is locked, the byte
>> range discloses where it is actively written, and the
>> appearance/disappearance of entries reflects host task activity,
>> contrary to the documented per-pidns visibility of /proc/locks
>> (proc_locks(5)).
>>
>> OFD locks record the owner tgid in flc_pid, so use it for the
>> visibility decision in locks_show(), mirroring what
>> locks_translate_pid() does for POSIX locks. The pid column is still
>> reported as -1 for OFD locks; only the filter decision changes.
>> Remote locks keep their negative flc_pid and stay visible as before.
>>
>> Verified with an OFD and a POSIX write lock held in the initial pid
>> namespace while a process in a fresh pid namespace reads /proc/locks:
>> the OFD entry is visible without this patch and hidden with it, the
>> POSIX entry is hidden in both cases.
>>
>> Signed-off-by: Tao Cui <cuitao@xxxxxxxxxx>
>>
>> ---
>>
>> Changes since v1: keep remote OFD locks (negative flc_pid) visible in
>> non-initial pid namespaces, as locks_translate_pid() does for remote
>> POSIX locks.
>> ---
>> fs/locks.c | 19 +++++++++++++++++++
>> 1 file changed, 19 insertions(+)
>>
>> diff --git a/fs/locks.c b/fs/locks.c
>> index 6e4ff7fcec05..4af1385682e4 100644
>> --- a/fs/locks.c
>> +++ b/fs/locks.c
>> @@ -3022,6 +3022,25 @@ static int locks_show(struct seq_file *f, void *v)
>>
>> cur = hlist_entry(v, struct file_lock_core, flc_link);
>>
>> + /*
>> + * OFD locks are reported with pid -1, so the filter below cannot see
>> + * their owner; flc_pid holds the owner tgid, so filter on it.
>> + * Remote locks keep a negative flc_pid and stay visible as before.
>> + */
>> + if ((cur->flc_flags & FL_OFDLCK) && cur->flc_pid > 0 &&
>> + proc_pidns != &init_pid_ns) {
>> + struct pid *pid;
>> + bool visible = false;
>> +
>> + rcu_read_lock();
>> + pid = find_pid_ns(cur->flc_pid, &init_pid_ns);
>> + if (pid)
>> + visible = pid_nr_ns(pid, proc_pidns) != 0;
>> + rcu_read_unlock();
>> + if (!visible)
>> + return 0;
>> + }
>> +
>> if (locks_translate_pid(cur, proc_pidns) == 0)
>> return 0;
>>
>
> First: I think this may be the wrong place to do this. Why not fold
> this change into locks_translate_pid()? It seems like we'd have
> inconsistent results wrt lock visibility between /proc/locks and
> F_GETLK if you do this here.
>

Good point. I did consider that. My hesitation with folding it into
locks_translate_pid() is that posix_lock_to_flock() also uses it, so
F_GETLK/F_OFD_GETLK would then report a conflicting OFD lock with
l_pid = 0 instead of the -1 that fcntl(2) documents for OFD locks.
That's why I kept the check in locks_show(), where it only affects
the /proc/locks view.

That distinction isn't introduced by this patch, though, since it
already exists for POSIX locks: since the 4.9 pid namespace filter, a
foreign-namespace POSIX lock that blocks F_GETLK is reported there
with l_pid = 0 while /proc/locks hides it. v2 just gives OFD locks
the same treatment.

> Thinking about this some more though, I wonder if trying to hide these
> locks is the right thing to do. These locks do exist and they do block
> you from acquiring a lock. If you go to look at /proc/locks and don't
> see them, that seems confusing.
>
> Would we be better off showing all the locks and reporting the pid as a
> negative value for ones acquired in foreign namespaces, like we do for
> remote locks?

That's a reasonable alternative - anonymizing the owner instead of
hiding the lock does have precedent (fdinfo follows that model), and I
can see the debugging benefit of always showing locks that may block
acquisition. I went with hiding because it preserves the /proc/locks
visibility semantics introduced in 4.9 by making OFD locks consistent
with POSIX locks. With only the pid anonymized, the output would
still disclose which files are locked and the locked byte ranges of
tasks outside the reader's pid namespace.

If the consensus is that /proc/locks should show all locks with an
anonymized pid instead, I'm happy to rework the patch in that
direction.

Thanks,
Tao