zhuxiangyi opened a new pull request, #10071:
URL: https://github.com/apache/paimon/pull/10071

   ### Purpose
   
   Bug fix: a data-evolution `MERGE INTO` whose source is a **different** table 
that happens to have the same `database.table` name in another catalog succeeds 
without doing anything — the source is silently ignored.
   
   #### Symptom
   
   ```sql
   -- paimon and paimon2 are two catalogs with different warehouses; each has a 
data-evolution table test.t
   -- paimon.test.t : (1, 10), (2, 20)      paimon2.test.t : (1, 100), (2, 200)
   MERGE INTO paimon.test.t AS t
   USING paimon2.test.t AS s
   ON t._ROW_ID = s._ROW_ID
   WHEN MATCHED THEN UPDATE SET v = s.v;
   
   SELECT id, v FROM paimon.test.t;   -- (1, 10), (2, 20): unchanged, and no 
error
   ```
   
   The statement returns normally, but `v` keeps its old values. A `WHEN 
MATCHED AND s.v > 100` condition is evaluated against the target's own `v` for 
the same reason, and adding `WHEN NOT MATCHED` trips the self-merge assertion 
("Self-Merge on _ROW_ID only supports WHEN MATCHED actions") even though this 
is not a self-merge at all.
   
   #### Impact
   
   - Any two Paimon catalogs whose tables share a `database.table` name: a 
staging/production pair with the same schema layout, a second catalog over 
another warehouse, a REST catalog next to a filesystem one. Copying a column 
from one into the other by row id is exactly the pattern that hits this.
   - The failure is silent: the MERGE reports success and the target is left as 
it was, so nothing points at the source having been ignored.
   - Not affected: branches (`t$branch_x` or the `branch` option both change 
the identifier), time travel on the source (already excluded), sources that are 
subqueries with filters or views (they do not match the passthrough shape), and 
same-named tables in another database (`database.table` differs).
   
   #### Root cause
   
   `MergeIntoPaimonDataEvolutionTable` has a self-merge shortcut for `MERGE 
INTO t USING t ON t._ROW_ID = s._ROW_ID`: it drops the source scan and rewrites 
every source attribute to the target's own attribute (`rewriteSourceToTarget`), 
so `SET v = s.v` becomes `SET v = t.v`. Whether the source *is* the target was 
decided by `sameSourceAndTargetTable`, which compared 
`DataSourceV2Relation.name`. For a Paimon table that is 
`Identifier.getFullName()` = `database.object` — the catalog is not part of it 
— so `paimon.test.t` and `paimon2.test.t` compare equal.
   
   #### Fix
   
   Decide table identity by what actually identifies a Paimon table: the 
storage location plus the (normalized) branch of the `FileStoreTable` behind 
each relation (`paimonTableIdentity`). Names no longer take part. Consequences:
   
   - a same-named table in another catalog, another branch of the same table, 
or a same-named table in another database → general join path, source values 
applied;
   - the same table reached through a second catalog that points at the same 
warehouse → still a self-merge, the shortcut is kept;
   - the unconditional/conditional `UPDATE` rewrite, which builds a self-merge 
from the same pinned `SparkTable`, is unaffected.
   
   The Spark 4.0 module keeps its own copy of this command; it is updated 
identically.
   
   ### Tests
   
   `RowTrackingTestBase` (run as `RowTrackingTest` on Spark 3.5, 70 tests, 
together with `BlobUpdateTest` and `DataEvolutionDeletionTest`):
   
   - `self-merge falls back for a same-named table in another catalog` — 
registers a second `SparkCatalog` over another warehouse, creates 
`paimon.test.t` and `paimon2.test.t` with different data, merges on `_ROW_ID` 
with a `WHEN MATCHED AND s.v > 100` condition and asserts the general plan (a 
`Join` is present), the target picks up the source values, the source is 
untouched, and a second MERGE with `WHEN NOT MATCHED ... INSERT` inserts the 
extra source row. **Fails on master**: the plan has no `Join` and the target 
keeps its old values.
   - `self-merge shortcut applies to the same table through another catalog` — 
a second catalog over the *same* warehouse; the MERGE still takes the shortcut 
(no `Join` / `Sort` / `RepartitionByExpression`) and `SET v = s.v + 1` 
increments.
   - `self-merge falls back for another branch of the same table` — creates a 
branch from a tag, changes the branch's values with `UPDATE`, merges from `` 
`t$branch_b1` `` and checks the join path is used and the branch values land in 
main.
   - `self-merge falls back for a same-named table in another database` — 
`test.t` vs `test2.t`, join path, source values applied.
   - The existing self-merge tests (`USING target`, aliases, passthrough 
projections, partition residuals, action-predicate pruning, pinned-snapshot 
`UPDATE`) keep passing, so the shortcut is still taken where it was before.
   
   `spotless:check` and `checkstyle:check` pass on both modules.
   
   ### API and Format
   
   No changes.
   
   ### Documentation
   
   No changes.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to