On 9/29/2026 9:07 PM, [email protected] wrote:
[You don't often get email from [email protected]. Learn why this is
important at https://aka.ms/LearnAboutSenderIdentification ]
Thank you for your contribution! Sashiko AI review found 1 potential issue(s)
to consider:
- [High] amdxdna: submitter thread hangs forever on failed device resume
--
commit 3c6d084a73c19b1b3a172be33a5eb12a9d7f8889
Author: David Zhang <[email protected]>
accel/amdxdna: Implement AIE4 suspend and resume
Implement suspend and resume callbacks for AIE4 Physical Function (PF),
Virtual Function (VF), and Classic device types.
diff --git a/drivers/accel/amdxdna/aie4_pci.c b/drivers/accel/amdxdna/aie4_pci.c
index f180983a692dc..6a50c1499ec93 100644
--- a/drivers/accel/amdxdna/aie4_pci.c
+++ b/drivers/accel/amdxdna/aie4_pci.c
[ ... ]
@@ -1096,11 +1096,250 @@ static void aie4_debugfs_init(struct amdxdna_dev *xdna)
[ ... ]
+static int aie4_vf_resume(struct amdxdna_dev *xdna)
+{
[ ... ]
+ ret = aie4_hwctx_resume_all(ndev);
+ if (ret) {
+ XDNA_ERR(xdna, "hwctx_resume failed %d", ret);
+ goto hw_clear;
+ }
+
+ XDNA_DBG(xdna, "vf resume done");
+ return 0;
+
+hw_clear:
+ aie4_hwctx_suspend_all(ndev, true);
+ aie4_vf_hw_stop(ndev);
+pci_disable:
+ pci_disable_device(pdev);
[Severity: High]
Could this error path leave submitter threads permanently hung in an
uninterruptible wait?
When aie4_vf_resume() or aie4_classic_resume() fails, they call
aie4_hwctx_suspend_all(ndev, true) which destroys the context with
AIE4_HWCTX_ERROR (setting has_reset = true).
However, a submitter thread waiting in wait_till_connected_hsa_not_full() in
drivers/accel/amdxdna/aie4_ctx.c with wait_through_reset = true explicitly
ignores has_reset:
ret = wait_event_freezable(priv->job_list_wq,
aie4_hwctx_connected(hwctx) ||
(!wait_through_reset &&
aie4_hwctx_has_reset(hwctx)));
Since the hardware resume failed, the context will never be recreated and
connected, causing it to block forever. This holds the srcu read lock and
prevents device removal (causing a deadlock in synchronize_srcu()).
I will fix the synchronize_srcu() deadlock in context teardown by stopping
the hardware context and waking all waiters before destroying it.
The fix will be in upcoming PATCH V2.
/David
+ return ret;
+}
--
Sashiko AI review ยท
https://sashiko.dev/#/patchset/[email protected]?part=16