yujun777 opened a new pull request, #68662: URL: https://github.com/apache/doris/pull/68662
### What problem does this PR solve? Related PR: #68390 Problem Summary: An `INSERT OVERWRITE` is two-phase: the rows are committed into temporary partitions first, and a later swap publishes them. A cancellation that lands between the two halves cannot take anything back -- the rows are durable, and everything the read consumed, the base table stream offsets among it, was committed with that same transaction (`InsertIntoTableCommand` hangs the offsets on the transaction, `DatabaseTransactionMgr.updateCatalogAfterCommitted` applies them on commit and replays them). The command answered that window by dropping the temporary partitions and returning normally: - the client was told the overwrite succeeded, while the table still held the rows it had; - the offsets had advanced past rows no target holds, so a re-run reads from the advanced offset and that range is silently missing. This is the general shape of `INSERT OVERWRITE t SELECT * FROM stream(...)`, and also of an IVM partition refresh, which resets the offsets of exactly the partitions it replaces. The cancellations that land before the rows are committed were answered the same way: success for a statement that did nothing. Fix: decide by whether the rows are already committed. - Before they are: nothing durable happened, so the statement fails -- like the cancellation the inner insert already reports for itself -- and a re-run reads the same rows. - After they are: the cancellation is too late, so the swap runs, the offset advance and the publication stay aligned, and the success the statement reports is the outcome it is. A debug point (`stage=beforeTheInsert|afterTheInsert`, scoped by table name) injects the cancellation at either side, in the same shape as the existing `failBetweenTheTwoHalvesOfAnOverwrite`. Scope: this closes the cancellation, which is the one entry into the window that reports success and therefore cannot be recovered by any later refresh. #68390's per-partition rebuild requirement, raised before a refresh reads, covers the failure and master-switch halves of the same window -- but a cancellation defeats it, precisely because the statement says it succeeded and the refresh records its epochs as met. The crash window and a `replacePartition` that throws are unchanged and still need the swap to become a committed action of the insert transaction. ### Release note `KILL`/cancellation of an `INSERT OVERWRITE` that lands after its rows were committed now completes the overwrite instead of silently reporting a success that published nothing; a cancellation before that point now reports an error instead of reporting success for a statement that did nothing. ### Check List (For Author) - Test: Regression test. New suite `insert_overwrite_p0/test_insert_overwrite_cancel` covers both sides; `insert_overwrite_p0` in full (14 suites) and `mtmv_p0/ivm/test_ivm_overwrite_failure_between_the_halves` pass locally. The new suite cannot fail on the pre-change code: the injection point it needs is introduced by this change, the same way #68390 added the point its suite uses. - Behavior changed: Yes. A cancelled `INSERT OVERWRITE` reports failure when it committed nothing, and success only after completing the swap, instead of reporting success in both cases. - Does this need documentation: No 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
