vc4->job_lock is a device-wide spinlock taken by the V3D interrupt handler,
so everything under it runs with interrupts disabled. It currently
protects several unrelated things: job queues, the binner slot pool, the
saved hang state, the active perfmon, the attachment of the job fence to
every BO, and more.

This series aims to (1) split vc4->job_lock into multiple independent locks
and (2) shrinks its critical sections when possible.

In particular, the series address four current issues:

  1. drm_syncobj_replace_fence() runs under job_lock, which lockdep reports
     as an IRQ lock inversion.

  2. drm_exec_fini() runs under it as well and reaches kvfree(), which asks
     for preemptible task context.

  3. VC4_SUBMIT_CL re-reads vc4->emit_seqno after dropping the lock, so a
     concurrent submission makes the ioctl return another job's seqno.

  4. The perfmon counters are copied to userspace with no lock held.

The rest of series does a series of changes to move work out of job_lock
and reduce contention.

On a Raspberry Pi 3B+ running `glmark2-wayland --off-screen`, lock_stat
shows job_lock contentions reduced by 90% (2093 -> 207 per run) and total
wait time by 96% (14.7 ms -> 0.6 ms). The longest job_lock hold, during
which interrupts are disabled, drops from 747 us to 27 us, as the render
done handler no longer signals fences under the lock. The workload is
GPU-bound, so the glmark2 score is unchanged (239.6 vs 240.2) [1].

About 99k acquisitions per run move off job_lock onto the per-fence locks,
and 185k more onto bin_alloc_lock, with little contention on either.

Considering that glmark2 is a GPU-bound client, it's important to notice
that contention improvements might be higher with several clients or a
CPU-bound workload, which I did not measure.

This series applies on top of "drm/vc4: Fix binner slot allocation failing
on an idle GPU" [2].

[1] `glmark2-wayland --off-screen` scores are the mean of 5 runs per kernel
without CONFIG_LOCK_STAT enabled. The lock data is the mean of 3 runs from
CONFIG_LOCK_STAT builds.

[2] 
https://lore.kernel.org/dri-devel/[email protected]/T/

Best regards,
- Maíra

---
Maíra Canal (11):
      drm/vc4: Publish the output fence outside of job_lock
      drm/vc4: Drop the BO reservations outside of job_lock
      drm/vc4: Refcount vc4_exec_info
      drm/vc4: Return the seqno of the submitted job
      drm/vc4: Attach the job fences outside of job_lock
      drm/vc4: Use the inline fence lock
      drm/vc4: Signal the render fence outside of job_lock
      drm/vc4: Take job_lock only once to drain job_done_list
      drm/vc4: Protect perfmon state with a dedicated lock
      drm/vc4: Protect the hang state with a dedicated lock
      drm/vc4: Split out bin_alloc_lock

 drivers/gpu/drm/vc4/vc4_drv.h     |  20 +++++--
 drivers/gpu/drm/vc4/vc4_gem.c     | 106 +++++++++++++++++++-------------------
 drivers/gpu/drm/vc4/vc4_irq.c     |  80 ++++++++++++++++------------
 drivers/gpu/drm/vc4/vc4_perfmon.c |  40 +++++++++-----
 drivers/gpu/drm/vc4/vc4_v3d.c     |   2 +-
 5 files changed, 146 insertions(+), 102 deletions(-)
---
base-commit: 64eded63d88ea70466c57cf433761bb1efea37a4
change-id: 20261002-vc4-misc-fixes-2b313622f570

Reply via email to