Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider: - [High] drm/panthor: stale pc->last_cmd state from untracked yield_idle() causes spurious command handling - [High] drm/panthor: synchronous state evaluation in INVALID_COMMAND handler races with hardware transitions - [High] drm/panthor: lockless state check in window_close() causes permanently leaked windows - [High] drm/panthor: overwriting pc->current_aw before CLOSE_WINDOW completes breaks state machine --
commit 137cc070e4a25e7d342b37a597b27e601b15068c Author: Karunika Choo <[email protected]> drm/panthor: Handle partition control INVALID_COMMAND interrupt This commit introduces handling for INVALID_COMMAND interrupts when partition state requirements are not met. It retries commands when the partition state permits, or requests scheduler recovery for open-window commands with inconsistent assignments. > diff --git a/drivers/gpu/drm/panthor/arbitration/panthor_partition_control.c > b/drivers/gpu/drm/panthor/arbitration/panthor_partition_control.c > index 95c26015cb637..446f77fc67891 100644 > --- a/drivers/gpu/drm/panthor/arbitration/panthor_partition_control.c > +++ b/drivers/gpu/drm/panthor/arbitration/panthor_partition_control.c [ ... ] > @@ -101,18 +114,45 @@ static int partition_state_wait(struct > panthor_partition_control *pc, u32 state) [ ... ] > static int yield_idle(struct panthor_partition_control *pc) > { > - gpu_write(pc->iomem, AM_PART_COMMAND, > - AM_PART_SET_COMMAND(PART_CMD_YIELD_IDLE)); > + guard(spinlock_irqsave)(&pc->lock); > + > + /* > + * This is a notification to yield the GPU once FW is idle. It is not > + * a command that must succeed. Nothing of note is lost if the command > + * fails apart from some minor scheduling optimisation losses. We can > + * skip tracking this command. > + */ > + partition_command_write_untracked( > + pc, AM_PART_SET_COMMAND(PART_CMD_YIELD_IDLE)); [Severity: High] Does sending PART_CMD_YIELD_IDLE untracked leave stale state in pc->last_cmd? If the scheduler calls yield_idle() and the hardware rejects the command, raising PART_INVALID_COMMAND, the interrupt handler will read the stale pc->last_cmd (for instance, from a previous OPEN_WINDOW) and act upon it. Can this lead to spurious command retries or fatal scheduler resets? > > return 0; > } > @@ -127,8 +167,8 @@ static int window_close(struct panthor_partition_control > *pc) > [Severity: High] Does the lockless hardware state check at the beginning of window_close() race with pending commands? If called immediately after window_open(), the hardware state may still be PART_STATE_RESET because the OPEN_WINDOW command is still in flight. window_close() assumes the window is already closed and returns early without sending CLOSE_WINDOW, leaving the window permanently open and leaking hardware resources. > pc->closing = true; > > - gpu_write(pc->iomem, AM_PART_COMMAND, > - AM_PART_SET_COMMAND(PART_CMD_CLOSE_WINDOW)); > + partition_command_write(pc, AM_PART_SET_COMMAND(PART_CMD_CLOSE_WINDOW), > + pc->current_aw); > > return 0; > } > @@ -153,8 +193,8 @@ static int window_open(struct panthor_partition_control > *pc, u8 aw_id) > > guard(spinlock_irqsave)(&pc->lock); > > - gpu_write(pc->iomem, AM_PART_COMMAND, > - AM_PART_SET_COMMAND(PART_CMD_OPEN_WINDOW) | > AM_PART_SET_WINDOW(aw_id)); > + partition_command_write(pc, (AM_PART_SET_COMMAND(PART_CMD_OPEN_WINDOW) | > + AM_PART_SET_WINDOW(aw_id)), aw_id); > > pc->current_aw = aw_id; [Severity: High] Does overwriting pc->current_aw before CLOSE_WINDOW completes break the state machine? In window_open(), if the window is already open, it calls window_close() and then immediately overwrites pc->current_aw with the new aw_id. When the hardware finishes closing and raises PART_RESET_DONE, the handler consumes the new current_aw, clears it to -1, and sends a spurious stop notification for the new window. Subsequent INVALID_COMMAND interrupts for the new window will be ignored. > > @@ -166,6 +206,8 @@ static void partition_handle_reset_done(struct > panthor_partition_control *pc) [ ... ] > +static void partition_handle_invalid_cmd(struct panthor_partition_control > *pc) > +{ [ ... ] > + switch (AM_PART_GET_COMMAND(pc->last_cmd.cmd)) { > + case PART_CMD_YIELD_NOW: > + case PART_CMD_CLOSE_WINDOW: > + if (partition_state_get(pc) == PART_STATE_WINDOW_OPEN) { > + partition_command_write_untracked(pc, > pc->last_cmd.cmd); > + return; > + } > + break; > + case PART_CMD_OPEN_WINDOW: > + if (partition_state_get(pc) == PART_STATE_RESET) { > + partition_command_write_untracked(pc, > pc->last_cmd.cmd); > + return; > + } else if (partition_aw_get(pc) == pc->last_cmd.aw_id) { > + return; > + } > + > + on_err = true; [Severity: High] Does this synchronous state evaluation race with hardware transitions? When OPEN_WINDOW is sent back-to-back after CLOSE_WINDOW, the hardware rejects it because it hasn't reached PART_STATE_RESET yet. The IRQ fires, and the handler synchronously reads the hardware state. If the hardware is still transitioning and hasn't reached RESET, the handler falls through to setting on_err = true, causing a fatal scheduler reset instead of retrying. > + aw_id = pc->last_cmd.aw_id; > + > + break; > + default: > + return; > + } > + } > + > + if (on_err) > + panthor_arbitration_notify_err(dev_get_drvdata(pc->dev), aw_id); > +} -- Sashiko AI review ยท https://sashiko.dev/#/patchset/[email protected]?part=20
