On 29/09/2026 11:54, Arunpravin Paneer Selvam wrote:
From: Arunpravin Paneer Selvam <[email protected]>

__alloc_range_bias() only undid splits made during its search when
split_block() itself failed. Its DFS-exhaustion failure path (-ENOSPC,
when no suitable block is found) skipped the undo, leaving the buddy
tree needlessly fragmented over repeated failed allocation attempts.

Fix by recording every successful split_block() call in a list and
unconditionally undoing those splits on every failure exit, via a
new single-level gpu_buddy_merge_one_level() helper (the original
__gpu_buddy_undo_splits() cascaded merges upward, which is unsafe
when called per split-list entry).

Resolves the igt@kms_plane@plane-panning-bottom-right@pipe-a/pipe-b
regression.

v2:
  - Drop the undo from __alloc_range(): it allocates every block it walks,
    so freeing that list on failure already merges the splits. (Matthew)

Fixes: 1ad5e807f716 ("gpu/buddy: replace dual-tree/force_merge with decoupled dirty 
tracker")
Assisted-by: Claude:claude-opus-4-8
Cc: Matthew Auld <[email protected]>
Cc: Christian König <[email protected]>
Signed-off-by: Arunpravin Paneer Selvam <[email protected]>
---
  drivers/gpu/buddy.c | 54 +++++++++++++++++++++++++++++++++++++++++++--
  1 file changed, 52 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/buddy.c b/drivers/gpu/buddy.c
index 2f2aaadafe35..e5c9e21cd077 100644
--- a/drivers/gpu/buddy.c
+++ b/drivers/gpu/buddy.c
@@ -1240,6 +1240,52 @@ static void __gpu_buddy_undo_splits(struct gpu_buddy *mm,
        }
  }
+static void gpu_buddy_merge_one_level(struct gpu_buddy *mm,
+                                     struct gpu_buddy_block *block)
+{
+       struct gpu_buddy_block *buddy = __get_buddy(block);
+       struct gpu_buddy_block *parent = block->parent;
+       enum gpu_block_state block_state;
+
+       if (!buddy || !gpu_buddy_block_is_free(block) ||
+           !gpu_buddy_block_is_free(buddy))
+               return;
+
+       block_state = gpu_block_cached_state(block);
+       if (gpu_block_cached_state(buddy) != block_state)
+               block_state = GPU_BLOCK_MIXED;
+
+       rbtree_remove(mm, block);
+       rbtree_remove(mm, buddy);
+       mm->free_scoreboard[gpu_buddy_block_order(block)] -= 2;
+
+       gpu_block_free(mm, block);
+       gpu_block_free(mm, buddy);
+
+       __mark_free(mm, parent, block_state);
+}
+
+static void gpu_buddy_undo_splits(struct gpu_buddy *mm,
+                                 struct gpu_buddy_block *block,
+                                 struct list_head *splits)
+{
+       if (block)
+               gpu_buddy_merge_one_level(mm, block);
+
+       while (!list_empty(splits)) {
+               struct gpu_buddy_block *parent =
+                       list_first_entry(splits, struct gpu_buddy_block,
+                                        tmp_link);
+
+               list_del(&parent->tmp_link);
+
+               if (!gpu_buddy_block_is_split(parent))
+                       continue;
+
+               gpu_buddy_merge_one_level(mm, parent->left);
+       }
+}
+
  static struct gpu_buddy_block *
  __alloc_range_bias(struct gpu_buddy *mm,
                   u64 start, u64 end,
@@ -1249,6 +1295,7 @@ __alloc_range_bias(struct gpu_buddy *mm,
        u64 req_size = mm->chunk_size << order;
        struct gpu_buddy_block *block;
        LIST_HEAD(dfs);
+       LIST_HEAD(splits);
        int err;
        int i;
@@ -1313,6 +1360,8 @@ __alloc_range_bias(struct gpu_buddy *mm,
                        err = split_block(mm, block);
                        if (unlikely(err))
                                goto err_undo;
+
+                       list_add(&block->tmp_link, &splits);

Looking at this now, I also don't see any issue here? If we got here:

1. block_order >= order.
2. adjust_end/start is aligned to order and fits within block.

Given that there must be something in here that will eventually hit block_order == order, given some number of splits, so either split fails or we must hit the 'return block'?


                }
/*
@@ -1349,7 +1398,7 @@ __alloc_range_bias(struct gpu_buddy *mm,
                }
        } while (1);
- return ERR_PTR(-ENOSPC);
+       err = -ENOSPC;
err_undo:
        /*
@@ -1357,7 +1406,8 @@ __alloc_range_bias(struct gpu_buddy *mm,
         * bigger is better, so make sure we merge everything back before we
         * free the allocated blocks.
         */
-       __gpu_buddy_undo_splits(mm, block);
+       gpu_buddy_undo_splits(mm, block, &splits);
+
        return ERR_PTR(err);
  }
base-commit: 90780f2c3d30187116128f71bcf92c8ab63400e7

Reply via email to