Thank you for your contribution! Sashiko AI review found 4 potential issue(s) 
to consider:
- [High] drm/panthor: stale pc->last_cmd state from untracked yield_idle() 
causes spurious command handling
- [High] drm/panthor: synchronous state evaluation in INVALID_COMMAND handler 
races with hardware transitions
- [High] drm/panthor: lockless state check in window_close() causes permanently 
leaked windows
- [High] drm/panthor: overwriting pc->current_aw before CLOSE_WINDOW completes 
breaks state machine
--

commit 137cc070e4a25e7d342b37a597b27e601b15068c
Author: Karunika Choo <[email protected]>

drm/panthor: Handle partition control INVALID_COMMAND interrupt

This commit introduces handling for INVALID_COMMAND interrupts when partition
state requirements are not met. It retries commands when the partition state
permits, or requests scheduler recovery for open-window commands with
inconsistent assignments.

> diff --git a/drivers/gpu/drm/panthor/arbitration/panthor_partition_control.c 
> b/drivers/gpu/drm/panthor/arbitration/panthor_partition_control.c
> index 95c26015cb637..446f77fc67891 100644
> --- a/drivers/gpu/drm/panthor/arbitration/panthor_partition_control.c
> +++ b/drivers/gpu/drm/panthor/arbitration/panthor_partition_control.c
[ ... ]
> @@ -101,18 +114,45 @@ static int partition_state_wait(struct 
> panthor_partition_control *pc, u32 state)
[ ... ]
>  static int yield_idle(struct panthor_partition_control *pc)
>  {
> -     gpu_write(pc->iomem, AM_PART_COMMAND,
> -               AM_PART_SET_COMMAND(PART_CMD_YIELD_IDLE));
> +     guard(spinlock_irqsave)(&pc->lock);
> +
> +     /*
> +      * This is a notification to yield the GPU once FW is idle. It is not
> +      * a command that must succeed. Nothing of note is lost if the command
> +      * fails apart from some minor scheduling optimisation losses. We can
> +      * skip tracking this command.
> +      */
> +     partition_command_write_untracked(
> +             pc, AM_PART_SET_COMMAND(PART_CMD_YIELD_IDLE));

[Severity: High]
Does sending PART_CMD_YIELD_IDLE untracked leave stale state in pc->last_cmd?

If the scheduler calls yield_idle() and the hardware rejects the command,
raising PART_INVALID_COMMAND, the interrupt handler will read the stale
pc->last_cmd (for instance, from a previous OPEN_WINDOW) and act upon it.
Can this lead to spurious command retries or fatal scheduler resets?

>  
>       return 0;
>  }
> @@ -127,8 +167,8 @@ static int window_close(struct panthor_partition_control 
> *pc)
>  

[Severity: High]
Does the lockless hardware state check at the beginning of window_close()
race with pending commands?

If called immediately after window_open(), the hardware state may still be
PART_STATE_RESET because the OPEN_WINDOW command is still in flight.
window_close() assumes the window is already closed and returns early
without sending CLOSE_WINDOW, leaving the window permanently open and
leaking hardware resources.

>       pc->closing = true;
>  
> -     gpu_write(pc->iomem, AM_PART_COMMAND,
> -               AM_PART_SET_COMMAND(PART_CMD_CLOSE_WINDOW));
> +     partition_command_write(pc, AM_PART_SET_COMMAND(PART_CMD_CLOSE_WINDOW),
> +                             pc->current_aw);
>  
>       return 0;
>  }
> @@ -153,8 +193,8 @@ static int window_open(struct panthor_partition_control 
> *pc, u8 aw_id)
>  
>       guard(spinlock_irqsave)(&pc->lock);
>  
> -     gpu_write(pc->iomem, AM_PART_COMMAND,
> -               AM_PART_SET_COMMAND(PART_CMD_OPEN_WINDOW) | 
> AM_PART_SET_WINDOW(aw_id));
> +     partition_command_write(pc, (AM_PART_SET_COMMAND(PART_CMD_OPEN_WINDOW) |
> +                                  AM_PART_SET_WINDOW(aw_id)), aw_id);
>  
>       pc->current_aw = aw_id;

[Severity: High]
Does overwriting pc->current_aw before CLOSE_WINDOW completes break the
state machine?

In window_open(), if the window is already open, it calls window_close()
and then immediately overwrites pc->current_aw with the new aw_id. When
the hardware finishes closing and raises PART_RESET_DONE, the handler
consumes the new current_aw, clears it to -1, and sends a spurious stop
notification for the new window. Subsequent INVALID_COMMAND interrupts
for the new window will be ignored.

>  
> @@ -166,6 +206,8 @@ static void partition_handle_reset_done(struct 
> panthor_partition_control *pc)
[ ... ]
> +static void partition_handle_invalid_cmd(struct panthor_partition_control 
> *pc)
> +{
[ ... ]
> +             switch (AM_PART_GET_COMMAND(pc->last_cmd.cmd)) {
> +             case PART_CMD_YIELD_NOW:
> +             case PART_CMD_CLOSE_WINDOW:
> +                     if (partition_state_get(pc) == PART_STATE_WINDOW_OPEN) {
> +                             partition_command_write_untracked(pc, 
> pc->last_cmd.cmd);
> +                             return;
> +                     }
> +                     break;
> +             case PART_CMD_OPEN_WINDOW:
> +                     if (partition_state_get(pc) == PART_STATE_RESET) {
> +                             partition_command_write_untracked(pc, 
> pc->last_cmd.cmd);
> +                             return;
> +                     } else if (partition_aw_get(pc) == pc->last_cmd.aw_id) {
> +                             return;
> +                     }
> +
> +                     on_err = true;

[Severity: High]
Does this synchronous state evaluation race with hardware transitions?

When OPEN_WINDOW is sent back-to-back after CLOSE_WINDOW, the hardware
rejects it because it hasn't reached PART_STATE_RESET yet. The IRQ fires,
and the handler synchronously reads the hardware state. If the hardware is
still transitioning and hasn't reached RESET, the handler falls through to
setting on_err = true, causing a fatal scheduler reset instead of retrying.

> +                     aw_id = pc->last_cmd.aw_id;
> +
> +                     break;
> +             default:
> +                     return;
> +             }
> +     }
> +
> +     if (on_err)
> +             panthor_arbitration_notify_err(dev_get_drvdata(pc->dev), aw_id);
> +}

-- 
Sashiko AI review ยท 
https://sashiko.dev/#/patchset/[email protected]?part=20

Reply via email to