zhuxiangyi opened a new pull request, #10071:
URL: https://github.com/apache/paimon/pull/10071
### Purpose
Bug fix: a data-evolution `MERGE INTO` whose source is a **different** table
that happens to have the same `database.table` name in another catalog succeeds
without doing anything — the source is silently ignored.
#### Symptom
```sql
-- paimon and paimon2 are two catalogs with different warehouses; each has a
data-evolution table test.t
-- paimon.test.t : (1, 10), (2, 20) paimon2.test.t : (1, 100), (2, 200)
MERGE INTO paimon.test.t AS t
USING paimon2.test.t AS s
ON t._ROW_ID = s._ROW_ID
WHEN MATCHED THEN UPDATE SET v = s.v;
SELECT id, v FROM paimon.test.t; -- (1, 10), (2, 20): unchanged, and no
error
```
The statement returns normally, but `v` keeps its old values. A `WHEN
MATCHED AND s.v > 100` condition is evaluated against the target's own `v` for
the same reason, and adding `WHEN NOT MATCHED` trips the self-merge assertion
("Self-Merge on _ROW_ID only supports WHEN MATCHED actions") even though this
is not a self-merge at all.
#### Impact
- Any two Paimon catalogs whose tables share a `database.table` name: a
staging/production pair with the same schema layout, a second catalog over
another warehouse, a REST catalog next to a filesystem one. Copying a column
from one into the other by row id is exactly the pattern that hits this.
- The failure is silent: the MERGE reports success and the target is left as
it was, so nothing points at the source having been ignored.
- Not affected: branches (`t$branch_x` or the `branch` option both change
the identifier), time travel on the source (already excluded), sources that are
subqueries with filters or views (they do not match the passthrough shape), and
same-named tables in another database (`database.table` differs).
#### Root cause
`MergeIntoPaimonDataEvolutionTable` has a self-merge shortcut for `MERGE
INTO t USING t ON t._ROW_ID = s._ROW_ID`: it drops the source scan and rewrites
every source attribute to the target's own attribute (`rewriteSourceToTarget`),
so `SET v = s.v` becomes `SET v = t.v`. Whether the source *is* the target was
decided by `sameSourceAndTargetTable`, which compared
`DataSourceV2Relation.name`. For a Paimon table that is
`Identifier.getFullName()` = `database.object` — the catalog is not part of it
— so `paimon.test.t` and `paimon2.test.t` compare equal.
#### Fix
Decide table identity by what actually identifies a Paimon table: the
storage location plus the (normalized) branch of the `FileStoreTable` behind
each relation (`paimonTableIdentity`). Names no longer take part. Consequences:
- a same-named table in another catalog, another branch of the same table,
or a same-named table in another database → general join path, source values
applied;
- the same table reached through a second catalog that points at the same
warehouse → still a self-merge, the shortcut is kept;
- the unconditional/conditional `UPDATE` rewrite, which builds a self-merge
from the same pinned `SparkTable`, is unaffected.
The Spark 4.0 module keeps its own copy of this command; it is updated
identically.
### Tests
`RowTrackingTestBase` (run as `RowTrackingTest` on Spark 3.5, 70 tests,
together with `BlobUpdateTest` and `DataEvolutionDeletionTest`):
- `self-merge falls back for a same-named table in another catalog` —
registers a second `SparkCatalog` over another warehouse, creates
`paimon.test.t` and `paimon2.test.t` with different data, merges on `_ROW_ID`
with a `WHEN MATCHED AND s.v > 100` condition and asserts the general plan (a
`Join` is present), the target picks up the source values, the source is
untouched, and a second MERGE with `WHEN NOT MATCHED ... INSERT` inserts the
extra source row. **Fails on master**: the plan has no `Join` and the target
keeps its old values.
- `self-merge shortcut applies to the same table through another catalog` —
a second catalog over the *same* warehouse; the MERGE still takes the shortcut
(no `Join` / `Sort` / `RepartitionByExpression`) and `SET v = s.v + 1`
increments.
- `self-merge falls back for another branch of the same table` — creates a
branch from a tag, changes the branch's values with `UPDATE`, merges from ``
`t$branch_b1` `` and checks the join path is used and the branch values land in
main.
- `self-merge falls back for a same-named table in another database` —
`test.t` vs `test2.t`, join path, source values applied.
- The existing self-merge tests (`USING target`, aliases, passthrough
projections, partition residuals, action-predicate pruning, pinned-snapshot
`UPDATE`) keep passing, so the shortcut is still taken where it was before.
`spotless:check` and `checkstyle:check` pass on both modules.
### API and Format
No changes.
### Documentation
No changes.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]