ashokraminedi opened a new pull request, #20121: URL: https://github.com/apache/hudi/pull/20121
### Describe the issue this Pull Request addresses Closes #19771 When Spark SQL `INSERT OVERWRITE TABLE` is executed using the row-writer bulk_insert path, the physical write operation is `BULK_INSERT` while the logical overwrite intent is carried through `BULKINSERT_OVERWRITE_OPERATION_TYPE` as `INSERT_OVERWRITE_TABLE`. During save-mode handling, only the physical operation was considered. As a result, SaveMode.Overwrite could treat the operation as a regular bulk insert and delete/re-initialize an existing table before the insert-overwrite-table executor was reached. This could remove existing timeline state, including pending clustering metadata, instead of allowing the overwrite-table path to handle the operation. ### Summary and Changelog This change preserves the existing table when row-writer bulk_insert is executing a logical `INSERT_OVERWRITE_TABLE`. Changes: - Detect `BULK_INSERT` operations whose `BULKINSERT_OVERWRITE_OPERATION_TYPE` is `INSERT_OVERWRITE_TABLE`. - Pass the logical `INSERT_OVERWRITE_TABLE` operation to save-mode handling so the existing table is not deleted and re-initialized. - Keep the physical write operation as `BULK_INSERT` for the subsequent row-writer execution path. - Add regression coverage for `INSERT OVERWRITE TABLE` with pending clustering. - Verify that the overwrite reaches the pending-clustering overlap check, existing table data remains intact when the operation is rejected, and the pending clustering instant is preserved. - Update the related test comment to reflect the corrected row-writer behavior. No code was copied. ### Impact No public API or configuration changes. The change is limited to Spark SQL row-writer bulk_insert handling for whole-table `INSERT OVERWRITE TABLE`. It prevents the existing table from being deleted during save-mode handling and allows the existing insert-overwrite-table execution path and conflict checks to run as intended. No expected performance impact. ### Risk Level low The change is narrowly scoped to `BULK_INSERT` when the internal overwrite operation type is explicitly `INSERT_OVERWRITE_TABLE`. Other bulk-insert operations continue using the existing save-mode behavior. Verification: - Added regression coverage for row-writer `INSERT OVERWRITE TABLE` with pending clustering. - Verified that the regression test reproduces the issue without the fix. - Ran `TestInsertTable2` after the fix: 22 tests passed, 0 failed. - Verified the existing table data and pending clustering instant remain intact when the overwrite is rejected. ### Documentation Update none No new feature, public API, or configuration is introduced, and no configuration defaults are changed. ### Contributor's checklist - [X] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [X] Enough context is provided in the sections above - [X] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
