On 29/09/2026 11:54, Arunpravin Paneer Selvam wrote:
From: Arunpravin Paneer Selvam <[email protected]>
__alloc_range_bias() only undid splits made during its search when
split_block() itself failed. Its DFS-exhaustion failure path (-ENOSPC,
when no suitable block is found) skipped the undo, leaving the buddy
tree needlessly fragmented over repeated failed allocation attempts.
Fix by recording every successful split_block() call in a list and
unconditionally undoing those splits on every failure exit, via a
new single-level gpu_buddy_merge_one_level() helper (the original
__gpu_buddy_undo_splits() cascaded merges upward, which is unsafe
when called per split-list entry).
Resolves the igt@kms_plane@plane-panning-bottom-right@pipe-a/pipe-b
regression.
v2:
- Drop the undo from __alloc_range(): it allocates every block it walks,
so freeing that list on failure already merges the splits. (Matthew)
Fixes: 1ad5e807f716 ("gpu/buddy: replace dual-tree/force_merge with decoupled dirty
tracker")
Assisted-by: Claude:claude-opus-4-8
Cc: Matthew Auld <[email protected]>
Cc: Christian König <[email protected]>
Signed-off-by: Arunpravin Paneer Selvam <[email protected]>
---
drivers/gpu/buddy.c | 54 +++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 52 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/buddy.c b/drivers/gpu/buddy.c
index 2f2aaadafe35..e5c9e21cd077 100644
--- a/drivers/gpu/buddy.c
+++ b/drivers/gpu/buddy.c
@@ -1240,6 +1240,52 @@ static void __gpu_buddy_undo_splits(struct gpu_buddy *mm,
}
}
+static void gpu_buddy_merge_one_level(struct gpu_buddy *mm,
+ struct gpu_buddy_block *block)
+{
+ struct gpu_buddy_block *buddy = __get_buddy(block);
+ struct gpu_buddy_block *parent = block->parent;
+ enum gpu_block_state block_state;
+
+ if (!buddy || !gpu_buddy_block_is_free(block) ||
+ !gpu_buddy_block_is_free(buddy))
+ return;
+
+ block_state = gpu_block_cached_state(block);
+ if (gpu_block_cached_state(buddy) != block_state)
+ block_state = GPU_BLOCK_MIXED;
+
+ rbtree_remove(mm, block);
+ rbtree_remove(mm, buddy);
+ mm->free_scoreboard[gpu_buddy_block_order(block)] -= 2;
+
+ gpu_block_free(mm, block);
+ gpu_block_free(mm, buddy);
+
+ __mark_free(mm, parent, block_state);
+}
+
+static void gpu_buddy_undo_splits(struct gpu_buddy *mm,
+ struct gpu_buddy_block *block,
+ struct list_head *splits)
+{
+ if (block)
+ gpu_buddy_merge_one_level(mm, block);
+
+ while (!list_empty(splits)) {
+ struct gpu_buddy_block *parent =
+ list_first_entry(splits, struct gpu_buddy_block,
+ tmp_link);
+
+ list_del(&parent->tmp_link);
+
+ if (!gpu_buddy_block_is_split(parent))
+ continue;
+
+ gpu_buddy_merge_one_level(mm, parent->left);
+ }
+}
+
static struct gpu_buddy_block *
__alloc_range_bias(struct gpu_buddy *mm,
u64 start, u64 end,
@@ -1249,6 +1295,7 @@ __alloc_range_bias(struct gpu_buddy *mm,
u64 req_size = mm->chunk_size << order;
struct gpu_buddy_block *block;
LIST_HEAD(dfs);
+ LIST_HEAD(splits);
int err;
int i;
@@ -1313,6 +1360,8 @@ __alloc_range_bias(struct gpu_buddy *mm,
err = split_block(mm, block);
if (unlikely(err))
goto err_undo;
+
+ list_add(&block->tmp_link, &splits);
Looking at this now, I also don't see any issue here? If we got here:
1. block_order >= order.
2. adjust_end/start is aligned to order and fits within block.
Given that there must be something in here that will eventually hit
block_order == order, given some number of splits, so either split fails
or we must hit the 'return block'?
}
/*
@@ -1349,7 +1398,7 @@ __alloc_range_bias(struct gpu_buddy *mm,
}
} while (1);
- return ERR_PTR(-ENOSPC);
+ err = -ENOSPC;
err_undo:
/*
@@ -1357,7 +1406,8 @@ __alloc_range_bias(struct gpu_buddy *mm,
* bigger is better, so make sure we merge everything back before we
* free the allocated blocks.
*/
- __gpu_buddy_undo_splits(mm, block);
+ gpu_buddy_undo_splits(mm, block, &splits);
+
return ERR_PTR(err);
}
base-commit: 90780f2c3d30187116128f71bcf92c8ab63400e7