zhuxiangyi opened a new pull request, #10391:
URL: https://github.com/apache/paimon/pull/10391
### Purpose
Bug fix: `sys.copy` leaves wrong row ids and sequence numbers in
row-tracking and data-evolution targets.
`sys.copy` (`CopyFilesCommitOperator`) commits the data files of the source
with the metadata they had there: `CopyFilesUtil.toNewDataFileMeta` keeps
`firstRowId`, `minSequenceNumber` and `maxSequenceNumber`. Three things go
wrong in the target:
1. **Rows written after the copy get the row ids of the copied rows.** A
row-tracking commit only advances `nextRowId` by the files it assigns row ids
to; files that arrive with a `firstRowId` do not count. After copying rows with
row ids `[0, 2)` into a new table, `nextRowId` stays 0 and the next `INSERT`
assigns 0 and 1 again: duplicate `_ROW_ID`s, and a global index or a `MERGE
INTO ... ON _ROW_ID` no longer identifies a row.
2. **A partial overwrite collides with the rows it keeps.** With
`dynamic-partition-overwrite` (the default), or a `where` filter, `sys.copy`
replaces only the partitions it copies into. The partitions it keeps have row
ids from the target, and the copied rows bring row ids from the source, which
can be the same.
3. **On a data-evolution target, later updates are ignored.** A
data-evolution read takes each column from the file with the highest sequence
number, and a commit stamps new files with its snapshot id. The copied files
keep the sequence numbers of the source, which can be far above the snapshot
ids of the target, so a later column update, `UPDATE` or `MERGE INTO` on the
target is silently hidden until the target's snapshot ids catch up.
This came up in the review of #10098 (converting an existing append table to
data evolution), where copied files with retained row ids led to rows hiding
each other.
### Changes
- `RowTrackingCommitUtils.assignRowTracking` (first commit): a commit that
adds files with a `firstRowId` they already have advances `nextRowId` past
them, and new files of the same commit get row ids after them.
- New `CopiedDataFiles.adaptToTarget`, called by `CopyFilesCommitOperator`
before the commit:
- **Row ids:** if the target has given out the copied row ids already (its
`nextRowId`, or the ranges of its live files), all copied row ids are shifted
by one offset past them. A single offset keeps the layout of files that share a
row id range, such as a data-evolution column update and the file it updates. A
copy into a table without rows keeps the row ids of the source.
- **Sequence numbers (data-evolution targets only):** the copied sequence
numbers, including the per-column ones of compacted files, are mapped in order:
the newest to 0, which the commit stamps with its snapshot id, older ones to
-1, -2, and so on. The copied files keep their order among themselves and are
all older than any later write to the target.
- **Files that store their row ids** (written by a copy-on-write `UPDATE`,
`DELETE` or `MERGE INTO` on a row-tracking table, no `firstRowId`): their row
ids are in their data, so they can neither be shifted nor covered by
`nextRowId`. Copying them into a row-tracking table is refused, asking to
rewrite those rows in the source first.
- Flink's copy clones the table metadata wholesale instead of committing
files, so it is not affected.
### Tests
- `RowTrackingCopiedFilesTest` (2): rows written after copied files get new
row ids; new files committed with copied files get row ids after them.
- `CopiedDataFilesTest` (12), copying the way `CopyFilesCommitOperator` does:
- row ids: copy into a table without rows keeps the source's row ids; copy
into one partition, and copy with a partition filter, shift past the rows of
the kept partitions; a copy over every row still shifts past the row ids given
out; a table without row tracking that already holds copied row ids; a shift
keeps a column update on its rows, which then take later updates by their new
row ids; files that store their row ids are refused by row-tracking targets,
also empty ones, and accepted by a table without row tracking;
- sequence numbers: later column updates win over copied sequence numbers;
the order of column updates from the source is kept; sequence numbers are
mapped in order, including per-column ones; row-tracking-only targets and
copies between tables without row tracking are unchanged.
- `CopyFilesProcedureTest` (Spark, real `CALL sys.copy`), 3 new: a copy with
a partition filter into a table with rows in another partition keeps all row
ids unique; after copying a data-evolution table, `UPDATE` and `MERGE INTO` on
the target take effect; a source file that stores its row ids is refused.
- Every change was checked by mutation: reverting it makes at least one of
these tests fail. The related core suites (`DataEvolution*`, `RowTracking*`,
`FileStoreCommitTest`, `AppendOnlySimpleTableTest`, `BlobTableTest`, 602 tests)
pass.
### API and Format
No API or format change. Behaviour change of `sys.copy` into row-tracking
tables: copied row ids may be shifted when the target has used them, and
copying files that store their row ids is refused.
### Documentation
No.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]