Gentle ping - anything needed from my side on this series?

Patchwork status so far: CI.Patch_applied, CI.Build and CI.KUnit all
green; CI.checkpatch has only the known warning for a quoted log line
in patch 2's commit message. The full hardware run does not seem to
have been triggered - could someone with permissions kick off the BAT
run on the series (patchwork series 172110)?

Meanwhile the two ARL test machines keep validating the fix: 17+
machine-days combined and >8M TLB invalidations since Wa_22016122933
was applied, still zero stalls - and the 7dd1 machine has now spent a
full week of normal use on xe (including plenty of video transcoding)
with zero platform freezes.

Thanks,
Tales


Em qua., 12 de ago. de 2026 às 23:30, Tales A. Mendonça
<[email protected]> escreveu:
>
> Hi,
>
> v2 of the TLB invalidation ack stall series for ARL, tracked in:
>
>   https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>
> The investigation concluded since v1: the GuC always replied in time.
> The CPU was reading a stale cache line of the G2H CTB - the GT1 caching
> bug that i915 works around as Wa_22016122933 and xe never inherited.
> Thanks Daniele for spotting the GT1 pattern in the GuC logs and
> suggesting the test, and Matt for the review and the guidance to mirror
> the i915 workaround.
>
> Patch 3 is new and is the actual fix: Wa_22016122933 for xe, scoped
> like i915 (media GT on media version 13.00) and covering the GuC-shared
> allocations the CPU reads from.
>
> Validation of patch 3 on two ARL machines (7d51 and 7dd1): before,
> 20-60 ack stalls/day, every day, for weeks, on two GuC firmware
> versions; after, zero stalls across 6+ days combined and >5M TLB
> invalidations. The 7dd1 machine, which could not survive a day of media
> workloads on xe without a platform freeze, ran 34h+ continuously
> including video transcoding with zero incidents.
>
> Patches 1-2 are the diagnostics that made the investigation possible,
> updated per Matt's v1 review.
>
> v1 -> v2:
> - Dropped the "kick" workaround (RFC patch 3): obsolete now that the
>   root cause is fixed by Wa_22016122933.
> - Patch 1: unified with the existing devcoredump path -
>   __xe_devcoredump() taking a GT plus nullable q/job, with
>   xe_devcoredump()/xe_devcoredump_gt() wrapper macros (Matt).
> - Patch 2: record the first timed out seqno instead of the last one,
>   and use xe_warn() (Matt).
> - New patch 3: implement Wa_22016122933.
>
> Based on today's drm-tip.
>
> Thanks,
> Tales
>
> Tales A. Mendonça (3):
>   drm/xe: Capture devcoredump on TLB invalidation timeout
>   drm/xe: Log when a timed out TLB invalidation ack finally arrives
>   drm/xe: Implement Wa_22016122933
>
>  drivers/gpu/drm/xe/xe_devcoredump.c     | 46 ++++++++++++++-----------
>  drivers/gpu/drm/xe/xe_devcoredump.h     | 15 +++++---
>  drivers/gpu/drm/xe/xe_guc.c             | 16 +++++++++
>  drivers/gpu/drm/xe/xe_guc.h             |  2 ++
>  drivers/gpu/drm/xe/xe_guc_ads.c         |  3 +-
>  drivers/gpu/drm/xe/xe_guc_ct.c          |  6 ++--
>  drivers/gpu/drm/xe/xe_guc_log.c         |  7 ++--
>  drivers/gpu/drm/xe/xe_guc_pc.c          |  3 +-
>  drivers/gpu/drm/xe/xe_tlb_inval.c       | 39 +++++++++++++++++++++
>  drivers/gpu/drm/xe/xe_tlb_inval_types.h | 17 +++++++++
>  drivers/gpu/drm/xe/xe_wa_oob.rules      |  1 +
>  11 files changed, 124 insertions(+), 31 deletions(-)
>
> --
> 2.55.0
>


-- 
Com os cumprimentos,

Tales A. Mendonça
talesam.org
communitybig.org

Reply via email to