zhuxiangyi commented on PR #10098: URL: https://github.com/apache/paimon/pull/10098#issuecomment-5993336527
@JingsongLi thanks for re-checking. All three points are answered in their threads: - **Copied row-id ranges** and **physical row-id files**: fixed in bbadbc65f, as you confirmed. 64aa7ce8c rebuilds the copied-file tests the way you asked: an ordinary target with copied rows only, copy 2 + append 2, and copy 2 + append 1. - **Sequence numbers of copied files on a plain target**: fixed in 8673cc17a. Every file from a schema without data evolution gets the baseline sequence 1. Your copy + column-update flow is a regression test in core and in Spark with the real `sys.copy`. **Separate PR for `sys.copy`: #10391.** Two of your findings come from `sys.copy` itself, not from the conversion. It commits the copied files with the `firstRowId` and sequence numbers of the source. That breaks row-tracking and data-evolution targets even without converting anything: 1. it does not advance `nextRowId`, so rows written after a copy get the copied rows' row ids again: duplicate `_ROW_ID`s; 2. a partial overwrite (dynamic partition overwrite, or a `where` filter) brings copied row ids that collide with the rows of the partitions it keeps; 3. on a data-evolution target, the copied sequence numbers can exceed the target's snapshot ids, so later `UPDATE` / `MERGE INTO` / column updates are silently ignored. #10391 fixes these in `sys.copy`: it advances `nextRowId`, shifts colliding copied row ids, maps the sequence numbers in order, and refuses files that store their row ids on row-tracking targets. This PR still has to handle copied files at conversion time, for two reasons: - your probes copy into an **ordinary** table, which keeps no row ids or `nextRowId`, so `sys.copy` has nothing to fix there; - tables copied before #10391 already hold such files. The two PRs are independent. The tests here build their state without relying on how `sys.copy` behaves, and they pass with and without #10391. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
