taoran92 opened a new issue, #10134: URL: https://github.com/apache/paimon/issues/10134
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar. ### Paimon version master ### Compute Engine Apache Spark 3.5.8, Scala 2.12. Reproduced with both V1 and V2 write configurations. ### Minimal reproduce step The following can be placed in a test in `DeleteFromTableTestBase`; `loadTable` is provided by the existing test base. Add imports for `org.apache.paimon.data.GenericRow` and `org.apache.paimon.types.RowKind`. ```scala spark.sql(""" CREATE TABLE T (id INT, v1 INT, seq1 INT, v2 INT, seq2 INT) TBLPROPERTIES ( 'primary-key' = 'id', 'bucket' = '1', 'merge-engine' = 'partial-update', 'fields.seq1.sequence-group' = 'v1', 'fields.seq2.sequence-group' = 'v2', 'partial-update.remove-record-on-sequence-group' = 'seq2', 'write-only' = 'true') """) // Two producers update independent field groups for the same keys. spark.sql("INSERT INTO T VALUES (1, 10, 1, NULL, NULL), (2, 20, 1, NULL, NULL)") spark.sql("INSERT INTO T VALUES (1, NULL, NULL, 100, 1), (2, NULL, NULL, 200, 1)") // Append a delete record with a newer seq2 for id=1. val builder = loadTable("T").newBatchWriteBuilder() val write = builder.newWrite() val commit = builder.newCommit() try { val delete = GenericRow.of(1, null, null, null, 2) delete.setRowKind(RowKind.DELETE) write.write(delete) commit.commit(write.prepareCommit()) } finally { write.close() commit.close() } spark.sql("SELECT * FROM T").show() spark.sql("SELECT COUNT(*) FROM T").show() spark.sql("SELECT id FROM T ORDER BY id").show() spark.sql("SELECT v1 FROM T ORDER BY v1").show() ``` `write-only=true` prevents write-time compaction so the query must merge the insert and delete records. The Paimon write API is used because SQL DELETE rewrites files for this configuration on this master revision, which would bypass the read-time merge scenario. ### What doesn't meet your expectations? The DELETE should remove the entire row for id=1 because its seq2=2 is newer than the stored seq2=1. Only `(2, 20, 1, 200, 1)` should remain, regardless of which columns are selected. | Query | Expected | Actual before the fix | | --- | --- | --- | | SELECT * FROM T | (2, 20, 1, 200, 1) | Correct | | SELECT COUNT(*) FROM T | 1 | 2 | | SELECT id FROM T ORDER BY id | 2 | 1, 2 | | SELECT v1 FROM T ORDER BY v1 | 20 | 10, 20 | Column pruning changes which rows are visible. The COUNT(*) execution plan contains an empty Paimon scan projection (`BatchScan test.T[]`). ### Anything else? `PartialUpdateMergeFunction.Factory.adjustReadType()` adds sequence dependencies for requested columns, but does not independently retain the sequence fields configured by `partial-update.remove-record-on-sequence-group`. These fields determine whether the whole row exists, even when none of that group's columns are requested. Consequently, pruning seq2 can prevent the merge function from recognizing the whole-row delete. The issue is in Paimon's core projection dependency handling and is triggered here by Spark column pruning; other compute engines have not been verified. A proposed fix is to include the whole-row deletion fields when deriving internal read dependencies and expand all comparator fields for composite sequence groups. The user-visible projection should remain unchanged. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
