Re: [PATCH v2 3/3] drm/panthor: Add tracepoints for cache flushing
From: Boris Brezillon
Date: Mon Aug 03 2026 - 05:15:04 EST
On Thu, 30 Jul 2026 13:45:16 +0200
Nicolas Frattaroli <nicolas.frattaroli@xxxxxxxxxxxxx> wrote:
> Add two new event tracepoints: gpu_cache_flush_start to be emitted after
> acquiring the flush mutex and reqs spinlock, and gpu_cache_flush_end to
> be emitted when leaving the function.
>
> This allows debugging the duration a flush takes irrespective of initial
> function entry lock contention by subtracting the start tracepoint's
> timestamp from the end tracepoint timestamp, and additionally contains
> information such as which caches were flushed.
>
> Signed-off-by: Nicolas Frattaroli <nicolas.frattaroli@xxxxxxxxxxxxx>
> ---
> drivers/gpu/drm/panthor/panthor_gpu.c | 3 ++
> drivers/gpu/drm/panthor/panthor_trace.h | 49 +++++++++++++++++++++++++++++++++
> 2 files changed, 52 insertions(+)
>
> diff --git a/drivers/gpu/drm/panthor/panthor_gpu.c b/drivers/gpu/drm/panthor/panthor_gpu.c
> index f015bde80abf..a25955ad668b 100644
> --- a/drivers/gpu/drm/panthor/panthor_gpu.c
> +++ b/drivers/gpu/drm/panthor/panthor_gpu.c
> @@ -336,6 +336,7 @@ int panthor_gpu_flush_caches(struct panthor_device *ptdev,
> guard(mutex)(&ptdev->gpu->cache_flush_lock);
>
> spin_lock(&ptdev->gpu->reqs_lock);
> + trace_gpu_cache_flush_start(ptdev->base.dev, l2, lsc, other);
> if (!(ptdev->gpu->pending_reqs & GPU_IRQ_CLEAN_CACHES_COMPLETED)) {
> ptdev->gpu->pending_reqs |= GPU_IRQ_CLEAN_CACHES_COMPLETED;
> gpu_write(gpu->iomem, GPU_CMD, GPU_FLUSH_CACHES(l2, lsc, other));
> @@ -345,6 +346,7 @@ int panthor_gpu_flush_caches(struct panthor_device *ptdev,
>
> if (ret) {
> spin_unlock(&ptdev->gpu->reqs_lock);
> + trace_gpu_cache_flush_end(ptdev->base.dev, l2, lsc, other);
Should we add a status to the end event, so that faulty flushes can
be filtered out (those can be immediate, or timeout=100ms depending on
where the failure happens, but they are not reflecting anything useful,
and would pollute the stats)? Actually, if what we care about is the
time it takes to do a flush, do we even need those start/end events,
can't we pass the time as an argument and do the math in the function
like we do in the job IRQ handler?
> return ret;
> }
>
> @@ -358,6 +360,7 @@ int panthor_gpu_flush_caches(struct panthor_device *ptdev,
> ptdev->gpu->pending_reqs &= ~GPU_IRQ_CLEAN_CACHES_COMPLETED;
> }
> spin_unlock(&ptdev->gpu->reqs_lock);
> + trace_gpu_cache_flush_end(ptdev->base.dev, l2, lsc, other);
>
> if (ret) {
> panthor_device_schedule_reset(ptdev);
> diff --git a/drivers/gpu/drm/panthor/panthor_trace.h b/drivers/gpu/drm/panthor/panthor_trace.h
> index 6ffeb4fe6599..6951b95b1de7 100644
> --- a/drivers/gpu/drm/panthor/panthor_trace.h
> +++ b/drivers/gpu/drm/panthor/panthor_trace.h
> @@ -76,6 +76,55 @@ TRACE_EVENT(gpu_job_irq,
> __entry->events, __entry->duration_ns)
> );
>
> +DECLARE_EVENT_CLASS(gpu_cache_flush_template,
> + TP_PROTO(const struct device *dev, u32 l2, u32 lsc, u32 other),
> + TP_ARGS(dev, l2, lsc, other),
> + TP_STRUCT__entry(
> + __string(dev_name, dev_name(dev))
> + __field(u32, l2)
> + __field(u32, lsc)
> + __field(u32, other)
> + ),
> + TP_fast_assign(
> + __assign_str(dev_name);
> + __entry->l2 = l2;
> + __entry->lsc = lsc;
> + __entry->other = other;
> + ),
> + TP_printk("%s: l2=0x%x lsc=0x%x other=0x%x", __get_str(dev_name),
> + __entry->l2, __entry->lsc, __entry->other)
> +);
> +
> +/**
> + * gpu_cache_flush_start - called after cache flush locks taken, before flush
> + * @dev: pointer to the &struct device, for printing the device name
> + * @l2: "l2" flush flags
> + * @lsc: "lsc" flush flags
> + * @other: "other" flush flags
> + *
> + * Fires after any initial lock contention around the locks needed for flushing
> + * caches, but before the actual cache flush is requested.
> + */
> +DEFINE_EVENT(gpu_cache_flush_template, gpu_cache_flush_start,
> + TP_PROTO(const struct device *dev, u32 l2, u32 lsc, u32 other),
> + TP_ARGS(dev, l2, lsc, other)
> +);
> +
> +/**
> + * gpu_cache_flush_end - called after cache flush
> + * @dev: pointer to the &struct device, for printing the device name
> + * @l2: "l2" flush flags
> + * @lsc: "lsc" flush flags
> + * @other: "other" flush flags
> + *
> + * Fires after either the cache flush is complete, or has failed. Can be used
> + * together with gpu_cache_flush_start to get how long the flush has taken.
> + */
> +DEFINE_EVENT(gpu_cache_flush_template, gpu_cache_flush_end,
> + TP_PROTO(const struct device *dev, u32 l2, u32 lsc, u32 other),
> + TP_ARGS(dev, l2, lsc, other)
> +);
> +
> #endif /* __PANTHOR_TRACE_H__ */
>
> #undef TRACE_INCLUDE_PATH
>