On 9/28/26 03:59, [email protected] wrote:
> From: Vitaly Prosyak <[email protected]>
> 
> The IGT amd_dispatch test exposed a NULL pointer dereference in the
> AMDGPU CS submission path when the GPU schedulers were not ready.
> 
> drm_sched_pick_best() returns NULL when every scheduler in an entity's
> list is marked not ready. drm_sched_entity_select_rq() then replaces
> the entity's existing runqueue with NULL.

That is perfectly correct behavior as far as I can see.

The schedulers should only be marked not ready when they are permanently dead 
and in this case keeping the existing rq doesn't make sense any more either.

> 
> A subsequent drm_sched_job_arm() retains that invalid runqueue.
> When AMDGPU CS submission calls drm_sched_entity_push_job(), the
> scheduler pointer derived from entity->rq is invalid and the access
> to sched->score faults. The reported oops shows the sequence:
> 
>     [drm] scheduler comp_1.1.0 is not ready, skipping
>     [drm] scheduler comp_1.2.0 is not ready, skipping
>     BUG: kernel NULL pointer dereference, address: 0000000000000268
>     RIP: drm_sched_entity_push_job+0x4f/0x2b0 [gpu_sched]
>     Call Trace:
>       amdgpu_cs_ioctl+0x1e9e/0x2530 [amdgpu]

That is clearly a bug in amdgpu. The scheduler behavior here is correct.

Regards,
Christian.

> 
> Keep the previously selected runqueue when no ready replacement is
> found. This prevents scheduler selection from turning a valid entity
> runqueue into NULL; it does not make a stopped scheduler ready or
> guarantee that the submitted job will execute.
> 
> Cc: Christian König <[email protected]>
> Cc: Alex Deucher <[email protected]>
> Cc: Matthew Brost <[email protected]>
> Cc: Danilo Krummrich <[email protected]>
> Cc: Philipp Stanner <[email protected]>
> Signed-off-by: Vitaly Prosyak <[email protected]>
> ---
>  drivers/gpu/drm/scheduler/sched_entity.c | 4 ++--
>  1 file changed, 2 insertions(+), 2 deletions(-)
> 
> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c 
> b/drivers/gpu/drm/scheduler/sched_entity.c
> index 4ebb513255ed..b11e1dddabd0 100644
> --- a/drivers/gpu/drm/scheduler/sched_entity.c
> +++ b/drivers/gpu/drm/scheduler/sched_entity.c
> @@ -584,8 +584,8 @@ void drm_sched_entity_select_rq(struct drm_sched_entity 
> *entity)
>  
>       spin_lock(&entity->lock);
>       sched = drm_sched_pick_best(entity->sched_list, entity->num_sched_list);
> -     rq = sched ? &sched->rq : NULL;
> -     if (rq != entity->rq) {
> +     if (sched && &sched->rq != entity->rq) {
> +             rq = &sched->rq;
>               drm_sched_rq_remove_entity(entity->rq, entity);
>               entity->rq = rq;
>       }

Reply via email to