Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
New issues: - [High] Returning 0 when the DAX layer returns `-EBUSY` silently clears the release event without notifying the device, causing a permanent capacity leak. Pre-existing issues: - [Critical] Unprotected `xa_for_each` iteration over `dc_extents` in `cxl_rm_extent` accesses unlocked and potentially freed `dc_extent` pointers, causing a Use-After-Free. -- commit a5dc5f1cb7a057f33cf2e7a95e0f09b9a990b420 Author: Ira Weiny <[email protected]> cxl + dax: Release dax_resources on DCD Release Capacity events Implement the release path that mirrors the add path: when the device asks for capacity back, the dax layer tears down the per-extent resources for the whole tag group atomically via dax_region_rm_resources(). > diff --git a/drivers/cxl/core/extent.c b/drivers/cxl/core/extent.c > index 59db1878b5e2d..7009ac6a51b46 100644 > --- a/drivers/cxl/core/extent.c > +++ b/drivers/cxl/core/extent.c [Severity: Critical] This isn't a bug introduced by this patch, but is it possible for the xa_for_each() loop in cxl_rm_extent() to trigger a use-after-free? Since cxl_rm_extent() and handle_add_event() hold cxl_rwsem.region only for read, can two threads execute them concurrently for the same region? If thread A calls rm_tag_group() and unregisters devices: drivers/cxl/core/extent.c:rm_tag_group() { xa_for_each(&group->dc_extents, index, dc_extent) devm_release_action(region_dev, dc_extent_unregister, dc_extent); } Which triggers a synchronous free via dc_extent_release(): drivers/cxl/core/extent.c:dc_extent_release() { kfree(dc_extent); } And concurrently, thread B executes the xa_for_each loop in cxl_rm_extent(): drivers/cxl/core/extent.c:cxl_rm_extent() { xa_for_each(&cxlr_dax->dc_extents, idx, dc_extent) { if (dc_extent->cxled != cxled) Since xa_for_each() takes and drops the RCU read lock internally during lookup, could thread B yield the dc_extent pointer just as thread A frees it, causing the access to dc_extent->cxled to be a use-after-free? [ ... ] > @@ -627,6 +627,18 @@ int cxl_rm_extent(struct cxl_memdev_state *mds, struct > cxl_extent *extent) > if (rc) > return rc; > > + rc = cxlr_notify_extent(cxlr, DCD_RELEASE_CAPACITY, group); > + if (rc) { > + /* > + * dax layer refused (-EBUSY) or failed (-ENOMEM, etc.). Do > + * not proceed to tear down the tag group — leave its > + * dax_resources alive so we do not free them out from under > + * live dev_dax ranges. The device will retry the release. > + */ > + return 0; > + } [Severity: High] Does returning 0 here when the DAX layer returns -EBUSY silently clear the release event and cause a permanent capacity leak? If cxlr_notify_extent() returns -EBUSY because capacity is actively mapped, returning 0 signals success to cxl_handle_dcd_event_records(). This will cause cxl_mem_get_records_log() to issue a CXL_MBOX_OP_CLEAR_EVENT_RECORD command, which clears the event from the device without ever sending the CXL_MBOX_OP_RELEASE_DC command. The comment mentions the device will retry the release, but section 8.2.9.2.1.3 of the CXL 3.1 specification states the device 'MAY reissue an event', not MUST. If the device does not actively reissue the event, will the host permanently lose track of the release request and leak the capacity? > + > + /* Release the entire tag group */ > rm_tag_group(group); > return 0; > } -- Sashiko AI review · https://sashiko.dev/#/patchset/[email protected]?part=22
