ashokraminedi opened a new pull request, #20121:
URL: https://github.com/apache/hudi/pull/20121

   ### Describe the issue this Pull Request addresses
   Closes #19771
   
   When Spark SQL `INSERT OVERWRITE TABLE` is executed using the row-writer 
bulk_insert path, the physical write operation is `BULK_INSERT` while the 
logical overwrite intent is carried through 
`BULKINSERT_OVERWRITE_OPERATION_TYPE` as `INSERT_OVERWRITE_TABLE`.
   
   During save-mode handling, only the physical operation was considered. As a 
result, SaveMode.Overwrite could treat the operation as a regular bulk insert 
and delete/re-initialize an existing table before the insert-overwrite-table 
executor was reached. This could remove existing timeline state, including 
pending clustering metadata, instead of allowing the overwrite-table path to 
handle the operation.
   
   ### Summary and Changelog
   
   This change preserves the existing table when row-writer bulk_insert is 
executing a logical `INSERT_OVERWRITE_TABLE`.
   
   Changes:
   
   - Detect `BULK_INSERT` operations whose 
`BULKINSERT_OVERWRITE_OPERATION_TYPE` is `INSERT_OVERWRITE_TABLE`.
   
   - Pass the logical `INSERT_OVERWRITE_TABLE` operation to save-mode handling 
so the existing table is not deleted and re-initialized.
   
   - Keep the physical write operation as `BULK_INSERT` for the subsequent 
row-writer execution path.
   
   - Add regression coverage for `INSERT OVERWRITE TABLE` with pending 
clustering.
   
   - Verify that the overwrite reaches the pending-clustering overlap check, 
existing table data remains intact when the operation is rejected, and the 
pending clustering instant is preserved.
   
   - Update the related test comment to reflect the corrected row-writer 
behavior.
   
   No code was copied.
   
   ### Impact
   
   No public API or configuration changes.
   
   The change is limited to Spark SQL row-writer bulk_insert handling for 
whole-table `INSERT OVERWRITE TABLE`. It prevents the existing table from being 
deleted during save-mode handling and allows the existing 
insert-overwrite-table execution path and conflict checks to run as intended.
   
   No expected performance impact.
   
   ### Risk Level
   
   low
   
   The change is narrowly scoped to `BULK_INSERT` when the internal overwrite 
operation type is explicitly `INSERT_OVERWRITE_TABLE`. Other bulk-insert 
operations continue using the existing save-mode behavior.
   
   Verification:
   
   - Added regression coverage for row-writer `INSERT OVERWRITE TABLE` with 
pending clustering.
   
   - Verified that the regression test reproduces the issue without the fix.
   
   - Ran `TestInsertTable2` after the fix: 22 tests passed, 0 failed.
   
   - Verified the existing table data and pending clustering instant remain 
intact when the overwrite is rejected.
   
   ### Documentation Update
   
   none
   
   No new feature, public API, or configuration is introduced, and no 
configuration defaults are changed.
   
   ### Contributor's checklist
   
   - [X] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [X] Enough context is provided in the sections above
   - [X] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to