Thank you for your contribution! Sashiko AI review found 6 potential issue(s) 
to consider:
- [High] drm/panthor: missing AW IRQ cleanup in panthor_device_init() error path
- [High] drm/panthor: background msg_retry_work runs concurrently with suspend
- [Medium] workqueue: 100% CPU busy loop in panthor_aw_msg_retry_work
- [Medium] drm/panthor: 5-second stall during device unplug due to disabled AW 
IRQs
- [Medium] drm/panthor: dropped AM messages cause 5-second timeout stalls
- [Low] drm/panthor: ignored -EOPNOTSUPP in AW handshake leads to unsupported 
protocol usage
--

commit 3a133f64881fd3095e93885e724b739d725e86af
Author: Karunika Choo <[email protected]>

drm/panthor: Add access-window support

Add the VM-side access-window component for v15 GPUs. The new code
handles AM message handshakes, requests GPU access from the arbiter,
waits for the window-open interrupt, and yields access on suspend.
Split panthor_hw_init into bind device and initialize gpu_info after
panthor_aw has ensured GPU access.

> diff --git a/drivers/gpu/drm/panthor/panthor_aw.c 
> b/drivers/gpu/drm/panthor/panthor_aw.c
> index 0000000000000..848fd64eb51b5
> --- /dev/null
> +++ b/drivers/gpu/drm/panthor/panthor_aw.c
[ ... ]
> @@ -118,17 +118,16 @@ static void panthor_aw_msg_retry_work(struct 
> work_struct *work)
>       struct panthor_aw *aw =
>               container_of(work, struct panthor_aw, msg_retry_work);
>       struct panthor_am_msg *msg = &aw->msg;
>       int ret;
>  
>       ret = panthor_am_msg_retry(msg);
>       if (ret == -EINVAL)
>               drm_warn(&aw->ptdev->base, "Send FIFO unexpectedly empty");
>  
>       if (ret == -EBUSY || ret == -EAGAIN)
>               queue_work(aw->wq, &aw->msg_retry_work);

[Severity: Medium]
Will this cause a busy loop pegging a CPU core? If panthor_am_msg_retry() 
returns -EBUSY (because the hardware AM_OUTGOING_MESSAGE FIFO is not empty), 
the workqueue immediately re-queues itself without any delay. Should this use
queue_delayed_work() or an interrupt-driven mechanism to back off while
hardware processes the request?

>  }
>  
>  static void panthor_aw_send_msg(struct panthor_aw *aw, u64 message)
>  {
>       struct panthor_device *ptdev = aw->ptdev;
>       int ret;
>  
>       ret = panthor_am_msg_send(&aw->msg, message);
>       if (ret == -ENOSPC)
>               drm_err(&ptdev->base, "Send FIFO is full");

[Severity: Medium]
Should we handle this error instead of silently swallowing it? If the FIFO is
full, the request is never buffered or dispatched, meaning callers will 
unnecessarily stall for 5 seconds waiting for a state transition response that
will never arrive.

>  
>       if (ret == -EBUSY) {
>               drm_dbg(&ptdev->base, "Pending messages, scheduling retry 
> work");
>               queue_work(aw->wq, &aw->msg_retry_work);
>       }
>  }
>  
>  static void panthor_aw_handshake_handle(struct panthor_aw *aw, u64 message)
>  {
>       struct panthor_device *ptdev = aw->ptdev;
>       bool acked = AM_MSG_ACK_GET(message);
>       u8 version = AM_MSG_VERSION_GET(message);
>       int ret;
>  
>       ret = panthor_am_msg_version_validate(&aw->msg, version);
>       if (ret == -EOPNOTSUPP)
>               drm_warn(&ptdev->base,
>                        "Msg protocol version less than minimum supported (%u 
> < %u)",
>                        version, AM_MSG_MIN_SUPPORTED_VERSION);

[Severity: Low]
Is it safe to ignore the -EOPNOTSUPP error and continue? Operating with an
unsupported firmware protocol version may cause undefined behavior or firmware
crashes. Should the code explicitly fail the initialization rather than
assuming it is safe to interact with an incompatible Arbiter version?

>  
>       if (!acked) {
>               u64 reply = VM_ARB_INIT_MAKE(1, aw->msg.version);
>  
>               panthor_aw_send_msg(aw, reply);
[ ... ]
> @@ -303,17 +302,15 @@ int panthor_aw_init(struct panthor_device *ptdev)
>  
>       ret = panthor_request_aw_irq(ptdev, &aw->irq, irq, WINDOW_IRQ_MASK,
>                                    aw->iomem + WINDOW_INT_BASE);
>       if (ret)
>               return ret;
>  
>       ret = panthor_aw_request(aw);
>       if (ret)
>               return ret;

[Severity: High]
Does this error path leak the active AW IRQ state? If panthor_aw_request() 
fails, we return without unwinding the active IRQ or pending work queue. 

In panthor_device_init(), if this function returns an error, we jump directly
to err_rpm_put, bypassing panthor_aw_unplug():

drivers/gpu/drm/panthor/panthor_device.c:panthor_device_init() {
    ...
        ret = panthor_aw_init(ptdev);
        if (ret)
                goto err_rpm_put;
    ...
}

Leaving a shared hardware IRQ unmasked while the device may be partially 
initialized or powered off could cause kernel panics, such as bus errors on 
MMIO access in the handler.

>  
>       return 0;
>  }
>  
>  void panthor_aw_unplug(struct panthor_device *ptdev)
>  {
>       struct panthor_aw *aw = ptdev->aw;
>  
>       if (!aw)
>               return;
>  
>       disable_work_sync(&aw->msg_retry_work);
>  
>       panthor_aw_irq_suspend(&aw->irq);

[Severity: Medium]
Will this sequence cause a 5-second stall during device unplug? 

In panthor_device_unplug(), panthor_aw_unplug() is called before
pm_runtime_put_sync_suspend():

drivers/gpu/drm/panthor/panthor_device.c:panthor_device_unplug() {
    ...
        panthor_aw_unplug(ptdev);

        pm_runtime_dont_use_autosuspend(ptdev->base.dev);
        pm_runtime_put_sync_suspend(ptdev->base.dev);
    ...
}

panthor_aw_unplug() disables and masks the AW IRQ. Then, the PM suspend path 
calls panthor_aw_suspend() which uses panthor_aw_state_wait() to wait for an 
interrupt-driven state transition to READY. Since the IRQ is masked before 
the suspend yield happens, the completion interrupt is never handled, 
guaranteeing a full 5-second timeout stall.

>  }
>  
>  int panthor_aw_resume(struct panthor_device *ptdev)
[ ... ]
> @@ -333,16 +332,15 @@ int panthor_aw_suspend(struct panthor_device *ptdev)
>  
>       if (atomic_read(&aw->state) == PANTHOR_AW_STATE_READY)
>               goto out_irq_suspend;
>  
>       if (panthor_aw_state_try_set(aw, PANTHOR_AW_STATE_GPU_GRANTED,
>               PANTHOR_AW_STATE_GPU_STOPPED))
>               ret = panthor_aw_yield(aw);
>       else
>               ret = panthor_aw_state_wait(
>                       aw, PANTHOR_AW_STATE_READY,
>                       PANTHOR_AW_STATE_TRANSITION_TIMEOUT_MS);
>  
>  out_irq_suspend:
>       panthor_aw_irq_suspend(&aw->irq);
>       return ret;

[Severity: High]
Does this exit path omit flushing or canceling the background work? If a
message is sent just before or during device suspend, msg_retry_work could be
queued. 

Unlike in panthor_aw_unplug(), panthor_aw_suspend() does not call 
disable_work_sync(&aw->msg_retry_work) before the device transitions to a 
suspended state. Since aw->wq is not allocated with WQ_FREEZABLE, the work
can run after the device's clocks and power domains are disabled, and
accessing powered-off MMIO registers would cause a bus error (kernel panic).

>  }

-- 
Sashiko AI review ยท 
https://sashiko.dev/#/patchset/[email protected]?part=22

Reply via email to