Re: [RFC v3 0/3] block: Introduce a BPF-based I/O scheduler
From: Kaitao Cheng
Date: Sat Oct 03 2026 - 07:16:11 EST
在 2026/10/3 17:36, Alexei Starovoitov 写道:
> On Sat, Oct 03, 2026 at 12:27 PM Kaitao Cheng <kaitao.cheng@xxxxxxxxx> wrote:
>> PFQ gives us a concrete policy to explore how well the UFQ interface
>> supports more involved scheduling decisions and to guide further work on
>> the framework. It has not yet been used in production, and further testing
>> and workload evaluation are needed.
>
> Third version in six months and still not a single number.
> v2 got replies from the bots only.
> Without a solid use case there is no point in polishing this.
Thank you very much for your review. This is a fairly large patch series,
and I have gone through several iterations because I was concerned that
my approach might not align with the community's expectations. I wanted
to get feedback early so that any architectural issues could be corrected
promptly. I am not seeking inclusion at this stage; I posted the series
for discussion. When I started developing this project, I wasn't entirely
sure it would work either, so I have been writing and testing as I go, haha!
I have actually run some fio tests locally. Initially, moving the I/O
scheduling policy into BPF caused a substantial performance regression.
After three iterations, however, the current implementation can match
the performance of native schedulers such as mq-deadline.
It also performs comparably to the none scheduler on slower SSDs, but
there is still a substantial performance gap on fast NVMe devices. My
analysis suggests that this is because the none scheduler can frequently
bypass ordering in the ctx queues and issue requests directly to the
driver (see blk_mq_try_issue_directly). I'll try to optimize it further.
I am also exploring and testing real-world use cases, which is why I
developed PFQ in the third patch. Next, I plan to evaluate PFQ with
foreground/background applications and with co-located containerized
online services and batch workloads. I may be able to share the results
in the next iteration.
>> In particular, I would appreciate suggestions on the boundary between the
>> UFQ framework and BPF policies, the struct_ops interface, and request
>> ownership and fallback handling.
>
> One global ufq_ops for all disks is not the best shape.
> I'd do it like bpf_qdisc.
Good suggestion. I'll give it a try. Thanks!
> The ownership is the bigger problem.
> The request sits in ctx->rq_lists and in a bpf map at the same time
> and the kernel relies on the prog to keep the two in sync.
> The prog holds rq->ref. rq holds q_usage_counter until
> __blk_mq_free_request(). One request that the prog didn't return
> from dispatch_req or left in a map and blk_mq_freeze_queue() waits
> forever.
> sched_ext has a watchdog that kicks the bpf scheduler out.
> Something like that is necessary here too.
When the BPF scheduler leaks an rq reference, system I/O does indeed stall.
However, because the requests are also kept in ctx->rq_lists, everything
returns to normal once the BPF scheduler exits. A watchdog that kicks the
BPF scheduler out is necessary, and it's also something I was planning to
implement. Thank you very much!
--
Thanks
Kaitao Cheng