Re: [PATCH] io_uring: do not charge the SQ/CQ rings to RLIMIT_MEMLOCK

From: Jens Axboe

Date: Wed Oct 07 2026 - 17:34:10 EST


On 10/7/26 2:17 PM, David Wei wrote:
> On 2026-10-06 16:59, Jens Axboe wrote:
>> On 10/6/26 6:57 AM, hengyul@xxxxxxxxxx wrote:
>>> From: Hengyu Liang <hengyul@xxxxxxxxxx>
>>>
>>> Commit 8078486e1d53 ("io_uring: use region api for SQ") and commit
>>> 81a4058e0cd0 ("io_uring: use region api for CQ") made io_uring_setup()
>>> allocate the rings with io_create_region().
>>>
>>> However, io_create_region() charges the memory to RLIMIT_MEMLOCK, and
>>> the rings had been exempt from that limit since commit 26bfa89e25f4
>>> ("io_uring: place ring SQ/CQ arrays under memcg memory limits"). As of
>>> now, a user without CAP_IPC_LOCK gets ENOMEM from io_uring_setup() when
>>> their rings exceed the limit, which is 8 MiB by default. PostgreSQL
>>> developers have already hit this in their io_uring tests [1].
>>>
>>> The issue can be reproduced with a simple liburing program, run as an
>>> unprivileged user:
>>>
>>>      #include <liburing.h>
>>>      #include <stdio.h>
>>>
>>>      int main(void)
>>>      {
>>>              static struct io_uring ring[64];
>>>              int i;
>>>
>>>              for (i = 0; i < 64; i++)
>>>                      if (io_uring_queue_init(4096, &ring[i], 0) < 0)
>>>                              break;
>>>              printf("%d rings\n", i);
>>>              return 0;
>>>      }
>>>
>>> Before those commits (v6.13), it prints "64 rings". After those commits
>>> (v6.14), it prints "21 rings".
>>>
>>> This patch makes io_create_region() take the user to charge, and passes
>>> no user for the SQ/CQ rings.
>>
>> Agree that this is a bug, stricter accounting may break use cases.
>> However, I think we can solve this simpler, and actually kill more code.
>> How about something like the below instead? Only apply accounting to
>> user backed memory, which is how it used to work too. Would be great if
>> you could take a look and also run your test case against it.
>
> Tested in a VM and confirmed that with the patch below the reproducer
> correctly allocates all 64 w/ a 8 MB RLIMIT_MEMLOCK.

Sending out a new set since I haven't heard back, and would be nice to get this fixed. Would be great if you could test those too.

> This change is really helpful for me as well, thank you for addressing
> this Hengyu.

Indeed!

--
Jens Axboe