Re: [RFC v6 3/3] drm/panthor: Create per queue priority workqueues

From: Boris Brezillon

Date: Mon Oct 05 2026 - 09:14:14 EST


Hello Tejun and Tvrtko,

On Fri, 02 Oct 2026 09:30:53 -1000
Tejun Heo <tj@xxxxxxxxxx> wrote:

> > + sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] = alloc_workqueue("panthor-drm", WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
> > + sched->submit_wq[PANTHOR_CSG_PRIORITY_LOW] = sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM];
> > + sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] = alloc_workqueue("panthor-drm-high", WQ_HIGHPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
> > + sched->submit_wq[PANTHOR_CSG_PRIORITY_RT] = alloc_workqueue("panthor-drm-rt", WQ_RT | WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
>
> For an unbound wq max_active applies to the whole wq, so this is two
> in-flight items per priority level across all the queues on the device,
> where the shared wq before had no effective limit. queue_run_job() blocks
> on sched->lock which tick_work() holds across FW round trips, so two
> blocked run_jobs would stall every other queue's run and free work at that
> level.

First off, panthor_sched::lock being a device-wide contention point for
submissions is something we plan to address (either by using a rw_lock
taken in read mode in the submit path and write mode in the scheduler
tick path, or by locking at a finer granularity).

> What's the reason for 2?

I think it was picked to keep the number of RT threads small, and
because we have this huge contention point, in ::run_job(), it was
deemed unimportant for now. Ultimately, if we were to size max_active
according to how fast the HW can dequeue, I guess we would go for
something like `max_active=panthor_sched::csg_slot_count`.

The other thing we need to address is the fact we now have proper
priority enforcement for submissions, but events are still processed
in a random order, so we probably want to have per-prio wq for OOM
handling, and we need a mechanism to process events in the CSG prio
order in panthor_sched_report_fw_events(). I'm also worried that an
RT-prio context would preempt the tick scheduled on
panthor_scheduler::wq, which is a regular-prio wq, thus making fairness
between RT contexts non-functional (if one RT context manages to queue
jobs fast enough, it would stay on the CSG slot with the highest prio
longer than we expect).

It's all these tiny details I'd like to sort out (or at least have a
plan for) before merging the per-prio-submit-wq stuff in panthor. This
being said, I don't think it should block the workqueue/drm_sched
changes, if they are deemed acceptable.

BTW, I apologize for being MIA for this long :-/.

Regards,

Boris