zhuxiangyi opened a new pull request, #10391:
URL: https://github.com/apache/paimon/pull/10391

   ### Purpose
   
   Bug fix: `sys.copy` leaves wrong row ids and sequence numbers in 
row-tracking and data-evolution targets.
   
   `sys.copy` (`CopyFilesCommitOperator`) commits the data files of the source 
with the metadata they had there: `CopyFilesUtil.toNewDataFileMeta` keeps 
`firstRowId`, `minSequenceNumber` and `maxSequenceNumber`. Three things go 
wrong in the target:
   
   1. **Rows written after the copy get the row ids of the copied rows.** A 
row-tracking commit only advances `nextRowId` by the files it assigns row ids 
to; files that arrive with a `firstRowId` do not count. After copying rows with 
row ids `[0, 2)` into a new table, `nextRowId` stays 0 and the next `INSERT` 
assigns 0 and 1 again: duplicate `_ROW_ID`s, and a global index or a `MERGE 
INTO ... ON _ROW_ID` no longer identifies a row.
   2. **A partial overwrite collides with the rows it keeps.** With 
`dynamic-partition-overwrite` (the default), or a `where` filter, `sys.copy` 
replaces only the partitions it copies into. The partitions it keeps have row 
ids from the target, and the copied rows bring row ids from the source, which 
can be the same.
   3. **On a data-evolution target, later updates are ignored.** A 
data-evolution read takes each column from the file with the highest sequence 
number, and a commit stamps new files with its snapshot id. The copied files 
keep the sequence numbers of the source, which can be far above the snapshot 
ids of the target, so a later column update, `UPDATE` or `MERGE INTO` on the 
target is silently hidden until the target's snapshot ids catch up.
   
   This came up in the review of #10098 (converting an existing append table to 
data evolution), where copied files with retained row ids led to rows hiding 
each other.
   
   ### Changes
   
   - `RowTrackingCommitUtils.assignRowTracking` (first commit): a commit that 
adds files with a `firstRowId` they already have advances `nextRowId` past 
them, and new files of the same commit get row ids after them.
   - New `CopiedDataFiles.adaptToTarget`, called by `CopyFilesCommitOperator` 
before the commit:
     - **Row ids:** if the target has given out the copied row ids already (its 
`nextRowId`, or the ranges of its live files), all copied row ids are shifted 
by one offset past them. A single offset keeps the layout of files that share a 
row id range, such as a data-evolution column update and the file it updates. A 
copy into a table without rows keeps the row ids of the source.
     - **Sequence numbers (data-evolution targets only):** the copied sequence 
numbers, including the per-column ones of compacted files, are mapped in order: 
the newest to 0, which the commit stamps with its snapshot id, older ones to 
-1, -2, and so on. The copied files keep their order among themselves and are 
all older than any later write to the target.
     - **Files that store their row ids** (written by a copy-on-write `UPDATE`, 
`DELETE` or `MERGE INTO` on a row-tracking table, no `firstRowId`): their row 
ids are in their data, so they can neither be shifted nor covered by 
`nextRowId`. Copying them into a row-tracking table is refused, asking to 
rewrite those rows in the source first.
   - Flink's copy clones the table metadata wholesale instead of committing 
files, so it is not affected.
   
   ### Tests
   
   - `RowTrackingCopiedFilesTest` (2): rows written after copied files get new 
row ids; new files committed with copied files get row ids after them.
   - `CopiedDataFilesTest` (12), copying the way `CopyFilesCommitOperator` does:
     - row ids: copy into a table without rows keeps the source's row ids; copy 
into one partition, and copy with a partition filter, shift past the rows of 
the kept partitions; a copy over every row still shifts past the row ids given 
out; a table without row tracking that already holds copied row ids; a shift 
keeps a column update on its rows, which then take later updates by their new 
row ids; files that store their row ids are refused by row-tracking targets, 
also empty ones, and accepted by a table without row tracking;
     - sequence numbers: later column updates win over copied sequence numbers; 
the order of column updates from the source is kept; sequence numbers are 
mapped in order, including per-column ones; row-tracking-only targets and 
copies between tables without row tracking are unchanged.
   - `CopyFilesProcedureTest` (Spark, real `CALL sys.copy`), 3 new: a copy with 
a partition filter into a table with rows in another partition keeps all row 
ids unique; after copying a data-evolution table, `UPDATE` and `MERGE INTO` on 
the target take effect; a source file that stores its row ids is refused.
   - Every change was checked by mutation: reverting it makes at least one of 
these tests fail. The related core suites (`DataEvolution*`, `RowTracking*`, 
`FileStoreCommitTest`, `AppendOnlySimpleTableTest`, `BlobTableTest`, 602 tests) 
pass.
   
   ### API and Format
   
   No API or format change. Behaviour change of `sys.copy` into row-tracking 
tables: copied row ids may be shifted when the target has used them, and 
copying files that store their row ids is refused.
   
   ### Documentation
   
   No.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to