taoran92 opened a new issue, #10134:
URL: https://github.com/apache/paimon/issues/10134

   ### Search before asking
   
   - [x] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   
   ### Paimon version
   
   master
   
   ### Compute Engine
   
   Apache Spark 3.5.8, Scala 2.12. Reproduced with both V1 and V2 write 
configurations.
   
   
   ### Minimal reproduce step
   
   The following can be placed in a test in `DeleteFromTableTestBase`; 
`loadTable` is provided by the existing test base. Add imports for 
`org.apache.paimon.data.GenericRow` and `org.apache.paimon.types.RowKind`.
   
   ```scala
   spark.sql("""
     CREATE TABLE T (id INT, v1 INT, seq1 INT, v2 INT, seq2 INT)
     TBLPROPERTIES (
       'primary-key' = 'id',
       'bucket' = '1',
       'merge-engine' = 'partial-update',
       'fields.seq1.sequence-group' = 'v1',
       'fields.seq2.sequence-group' = 'v2',
       'partial-update.remove-record-on-sequence-group' = 'seq2',
       'write-only' = 'true')
   """)
   
   // Two producers update independent field groups for the same keys.
   spark.sql("INSERT INTO T VALUES (1, 10, 1, NULL, NULL), (2, 20, 1, NULL, 
NULL)")
   spark.sql("INSERT INTO T VALUES (1, NULL, NULL, 100, 1), (2, NULL, NULL, 
200, 1)")
   
   // Append a delete record with a newer seq2 for id=1.
   val builder = loadTable("T").newBatchWriteBuilder()
   val write = builder.newWrite()
   val commit = builder.newCommit()
   try {
     val delete = GenericRow.of(1, null, null, null, 2)
     delete.setRowKind(RowKind.DELETE)
     write.write(delete)
     commit.commit(write.prepareCommit())
   } finally {
     write.close()
     commit.close()
   }
   
   spark.sql("SELECT * FROM T").show()
   spark.sql("SELECT COUNT(*) FROM T").show()
   spark.sql("SELECT id FROM T ORDER BY id").show()
   spark.sql("SELECT v1 FROM T ORDER BY v1").show()
   ```
   
   `write-only=true` prevents write-time compaction so the query must merge the 
insert and delete records. The Paimon write API is used because SQL DELETE 
rewrites files for this configuration on this master revision, which would 
bypass the read-time merge scenario.
   
   ### What doesn't meet your expectations?
   
   The DELETE should remove the entire row for id=1 because its seq2=2 is newer 
than the stored seq2=1. Only `(2, 20, 1, 200, 1)` should remain, regardless of 
which columns are selected.
   
   | Query | Expected | Actual before the fix |
   | --- | --- | --- |
   | SELECT * FROM T | (2, 20, 1, 200, 1) | Correct |
   | SELECT COUNT(*) FROM T | 1 | 2 |
   | SELECT id FROM T ORDER BY id | 2 | 1, 2 |
   | SELECT v1 FROM T ORDER BY v1 | 20 | 10, 20 |
   
   Column pruning changes which rows are visible. The COUNT(*) execution plan 
contains an empty Paimon scan projection (`BatchScan test.T[]`).
   
   ### Anything else?
   
   `PartialUpdateMergeFunction.Factory.adjustReadType()` adds sequence 
dependencies for requested columns, but does not independently retain the 
sequence fields configured by `partial-update.remove-record-on-sequence-group`. 
These fields determine whether the whole row exists, even when none of that 
group's columns are requested.
   
   Consequently, pruning seq2 can prevent the merge function from recognizing 
the whole-row delete. The issue is in Paimon's core projection dependency 
handling and is triggered here by Spark column pruning; other compute engines 
have not been verified.
   
   A proposed fix is to include the whole-row deletion fields when deriving 
internal read dependencies and expand all comparator fields for composite 
sequence groups. The user-visible projection should remain unchanged.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to