Daniele, Matt - now that rc1 is out, could you weigh in on the two open
questions from the cover letter? They are the only things I am unsure
about, and I would rather follow your preference than guess:

1. CPU mapping. This series uses XE_BO_FLAG_NEEDS_UC, which maps both
   the CPU side (uncached) and the GGTT side (XE_CACHE_NONE). i915
   documents the workaround as WC on the CPU side and UC on the GPU
   side. A CPU-WC variant with the GGTT side left at XE_CACHE_NONE was
   measured to be equally effective, but expressing it needs either a
   new BO flag or decoupling the GGTT cache-mode selection from
   XE_BO_FLAG_NEEDS_UC - swapping in XE_BO_FLAG_FORCE_WC alone
   silently relaxes the GPU side back to WB and the stalls return.
   Keep the stricter UC, or should I add the plumbing for i915 parity?

2. Scope. i915 applies the workaround to all media-GT shared memory,
   including LRC/ring state; this series covers the GuC-shared
   allocations (CTBs, log, ADS, SLPC, engine activity). Should xe match
   i915's full scope?

Either answer is a small change; I just do not want to pick the wrong
one unilaterally.

Meanwhile both test machines have moved to 7.2 with this exact series
applied (previously they ran a cruder CTB-only variant), and the
workaround shows up where it should - present on the media GT's OOB
list, absent on the primary GT:

  gt0/workarounds: (no 22016122933)
  gt1/workarounds: 22016122933

Still zero ack stalls on either machine. The 7dd1 machine, which used
to freeze under media workloads on xe, has been running xe full time
since 11 August, including a day with 16 video transcoding sessions,
with no stalls and no freezes.

Thanks,
Tales

Em seg., 24 de ago. de 2026 às 11:07, Tales A. Mendonça
<[email protected]> escreveu:
>
> Hi,
>
> v3 of the TLB invalidation ack stall fix for ARL, tracked in:
>
>   https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>
> Root cause recap: the GuC replies in time, but on the standalone media
> GT the CPU reads a stale cache line of the G2H CTB - the caching
> erratum i915 works around as Wa_22016122933 and xe never inherited.
> Patch 3 implements it, scoped like i915.
>
> Since v2, Navon John Lukose A/B tested the series on an ARL 7d51 and
> produced three results that are folded into this version:
>
>  * Making the mapping coherent instead of uncached does not work: a
>    GGTT PAT entry repurposed to WB|COH_2WAY still stalls
>    (request-to-ack 2290ms). Now documented in patch 3 - it is the
>    obvious question to ask about this patch, and the answer is
>    measured rather than assumed.
>
>  * xe_guc_engine_activity.c allocates GuC-written, CPU-read buffers on
>    the same GT without the workaround flags; the symptom there would
>    be stale busyness stats rather than hangs. Now covered.
>
>  * CPU-WC (matching i915's wording) is equally effective, but
>    XE_BO_FLAG_NEEDS_UC currently drives both the CPU mapping and the
>    GGTT cache mode, and naively switching to XE_BO_FLAG_FORCE_WC
>    relaxes the GPU side back to WB and brings the stalls back
>    (3 hits/20min measured).
>
> His Tested-by on patch 3 is from v2; the only code change since is the
> added engine activity coverage.
>
> Two questions I would rather have answered than decide unilaterally:
>
>  1. CPU mapping: keep the stricter UC as in this series, or add a way
>     to express i915's exact combination (CPU WC + GGTT XE_CACHE_NONE)?
>     The latter needs either a new BO flag or decoupling the GGTT
>     cache-mode selection from XE_BO_FLAG_NEEDS_UC. No measurable
>     throughput difference between the two.
>
>  2. Scope: i915 applies the workaround to all media-GT shared memory,
>     including LRC/ring state; this series covers the GuC-shared
>     allocations only (CTBs, log, ADS, SLPC, engine activity). Should
>     xe match i915's full scope?
>
> I also left out Fixes:/Cc: stable, since MTL/ARL is still
> require_force_probe in xe - happy to add them if you would prefer the
> fix to reach stable kernels.
>
> Validation of patch 3: two ARL machines (7d51 and 7dd1), ~4 weeks of
> combined runtime, >7M TLB invalidations, zero stalls. Before the fix
> both machines reproduced 20-60 stalls/day, every day, on every kernel
> and on two GuC firmware versions. The 7dd1 machine, which could not
> survive a day of media workloads on xe without a platform freeze, has
> been running xe continuously for two weeks including video
> transcoding, with zero incidents.
>
> Patches 1-2 are the diagnostics that made the investigation possible,
> unchanged from v2 except for a commit-message line wrap.
>
> v2 -> v3:
> - Rebased on drm-tip.
> - Patch 3: cover xe_guc_engine_activity.c; document the coherency
>   experiment and the UC-vs-WC tradeoff; add Suggested-by/Tested-by.
> - Patch 2: wrap a quoted log line to silence checkpatch.
> - checkpatch is clean on all three patches now.
>
> Thanks,
> Tales
>
> Tales A. Mendonça (3):
>   drm/xe: Capture devcoredump on TLB invalidation timeout
>   drm/xe: Log when a timed out TLB invalidation ack finally arrives
>   drm/xe: Implement Wa_22016122933
>
>  drivers/gpu/drm/xe/xe_devcoredump.c         | 46 +++++++++++----------
>  drivers/gpu/drm/xe/xe_devcoredump.h         | 15 +++++--
>  drivers/gpu/drm/xe/xe_guc.c                 | 16 +++++++
>  drivers/gpu/drm/xe/xe_guc.h                 |  2 +
>  drivers/gpu/drm/xe/xe_guc_ads.c             |  3 +-
>  drivers/gpu/drm/xe/xe_guc_ct.c              |  6 ++-
>  drivers/gpu/drm/xe/xe_guc_engine_activity.c |  6 ++-
>  drivers/gpu/drm/xe/xe_guc_log.c             |  7 +++-
>  drivers/gpu/drm/xe/xe_guc_pc.c              |  3 +-
>  drivers/gpu/drm/xe/xe_tlb_inval.c           | 39 +++++++++++++++++
>  drivers/gpu/drm/xe/xe_tlb_inval_types.h     | 17 ++++++++
>  drivers/gpu/drm/xe/xe_wa_oob.rules          |  1 +
>  12 files changed, 128 insertions(+), 33 deletions(-)
>
> --
> 2.55.0
>


-- 
Com os cumprimentos,

Tales A. Mendonça
talesam.org
communitybig.org

Reply via email to