On 8/4/2026 2:33 PM, Summers, Stuart wrote:
On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
Hi,
This series is a follow-up to the TLB invalidation ack stall I
have
been debugging on ARL, tracked in:
https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
and
7dd1 machines here, plus an independent Arc Pro 130T report on
the
issue above), TLB invalidation acks intermittently stall for
~2.3s.
The H2G request is consumed from the CTB immediately and the G2H
CTB
is empty the whole time - the firmware simply does not send the
ack
until much later. The fence timeout fires at 2.25s and the ack
lands
tens of ms after it. Userspace blocked on the invalidation
(compositor
buffer unmaps etc.) hitches for the full window.
Firstly, thanks for the patch!
I haven't looked in to all the details of the sighting you were
debugging, but we have had similar issues that were fixed in a
later
GuC version. I think around 70.60.0? It might be worth trying on
something later than that to see if that helps... (+Daniele)
I think this would require an AR on our end to make a new firmware
version available.
The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
aliases to MTL for firmware). (+Julia too).
Presumably, the GuC changelogs should indicate whether an issue
related
this has been fixed. If so, we need to update all GuC versions across
both i915 and Xe.
Right... I guess I'd still like to see if we can test this in GuC (or
get confirmation we can't for some reason) before committing something.
My worry is we will prevent bug reports like this by working around it
and miss critical bugs that need to be fixed in the right component.
[1]
https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
no
Is there a reason we don't just re-use the main xe_devcoredump()?
This is my suggestion: the main devcoredump infrastructure is job-
based,
so it cannot be used for hangs that are not associated with a job.
In my opinion, this is a gap on our end. Introducing something like
`xe_devcoredump_gt()`, which can be used for non-job-based hangs
(e.g.,
TLB invalidation timeouts like those addressed in this series, or
more
generally any GuC protocol hang), makes sense to me.
Ok makes sense. We can do that here. It would be nice to have a more
inclusive implementation that lets us call this from anywhere so we
aren't duplicating things around for different use cases. But not a
blocker here.
I haven't looked at the patch yet, but at a high level, adding
`xe_devcoredump_gt()` seems like a reasonable approach.
exec queue or job to blame - leaves a devcoredump with the GuC
log
and
CT state behind (Matt suggested capturing devcoredumps when we
discussed the issue; devcoredumps from both machines are attached
to
the issue above).
Patch 2 logs when the ack for a timed out invalidation finally
arrives. This is what established that the acks are late rather
than
lost.
Patch 3 is the RFC part: a delayed work that pokes the GuC
(status
register read, CT flush, doorbell ring) every 250ms while an ack
is
overdue. On my machines this converts the guaranteed 2.3s stall
into
I'm a little worried we're just papering over something here that
needs
to be addressed in GuC, particularly around GT going to sleep or
something around the time we're expecting a response, so the pings
on
registers might be prematurely waking things up which is something
we'd
want to happen in GuC, not the KMD.
In general, I agree with this. We should avoid papering over the
issue
and instead fix it properly in the GuC. That said, this workaround
provides a pretty strong data point, since it appears to get the TLB
invalidation unstuck.
So if we hit this issue I guess we're already going to have some
performance degredation and the workaround makes that better. I need to
look at the implementation, but we could be potentially introducing
performance penalties in other areas doing these pings.
Again, I'd like to see if we can fix this in the right place before
implementing a workaround for it. Hopefully Daniele or Julia can give
some direction there.
Are we seeing this on i915 at all? Given that Xe does not officially
support MTL/ARL and is missing several critical WAs for those platforms,
the approach so far has been to only update the GuC FW if it is required
for i915.
Looking at the GuC release notes, there have been a couple of
TLB-related fixes after 70.53, but they're both marked as only affecting
PVC and Xe2+ platforms, so no fixes seem to be available for ARL (or at
least they're not listed in the release notes).
Daniele
Thanks,
Stuart
Matt
Thanks,
Stuart
a
sub-500ms hiccup for the majority of occurrences; a minority of
severe
episodes ignore 8-9 consecutive doorbells, which points at the
GuC
firmware being internally blocked for the whole window. Full data
on
the issue. I am happy to rework the approach (different delay,
tying it to the G2H handler, dropping the status read, etc.) -
mainly
I would like the firmware side investigated, since no host-side
poke
can fix the severe cases.
Based on drm-tip. Tested for several days on both ARL machines
under
desktop and VM-heavy workloads.
Thanks,
Tales
Tales A. Mendonça (3):
drm/xe: Capture devcoredump on TLB invalidation timeout
drm/xe: Log when a timed out TLB invalidation ack finally
arrives
drm/xe: Kick GuC while TLB invalidation acks are overdue
drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
drivers/gpu/drm/xe/xe_tlb_inval.c | 131
+++++++++++++++++++++++-
drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
4 files changed, 243 insertions(+), 4 deletions(-)