Re: [PATCH v2 0/3] block: skip the blkcg walk in blk_cgroup_congested() when nothing is throttled

From: Tejun Heo

Date: Fri Aug 14 2026 - 13:02:18 EST


On Fri, Aug 14, 2026 at 09:56:36AM -0700, Usama Arif wrote:
> blk_cgroup_congested() walks the current task's blkcg ancestor chain on every
> readahead decision and, once swap is in use, on every anonymous and shmem
> folio allocation. The answer is almost always "no", but finding that out
> costs two loads per level on two cold cache lines, plus an out-of-line
> kthread_blkcg() and an RCU read-side pair. On a fleet profile of hosts
> running containers with 5-10 level hierarchies it costs about as much as all
> of mutex_lock(), 99.4% of it under __folio_throttle_swaprate().
>
> Patch 3 gates the walk on a global count of blkcgs with a non-zero
> congestion_count, so the common case is a load and a predicted branch.
>
> That only works if the count is correctly maintained, currently two teardown
> paths can leave a blkcg permanently marked congested. Today that only hurts
> tasks in the affected cgroup, but it hurts them for the life of the cgroup -
> readahead cut to a single page, async readahead skipped, and a throttle
> scheduled on every anonymous folio allocation. With a global gate it would
> cost every other task on the machine the walk as well. Patches 1 and 2 fix
> those two paths and stand on their own as bugfixes; patch 3 depends on them.

For the series,

Acked-by: Tejun Heo <tj@xxxxxxxxxx>

Thanks.

--
tejun