Akash3121 opened a new issue, #9993:
URL: https://github.com/apache/paimon/issues/9993

   ### Search before asking
   
   - [x] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   
   ### Paimon version
   
   master (1d368b4a5932f8221fd28e2555001abdb8fb12ee)
   
   
   ### Compute Engine
   
   Spark with a path-based Paimon table. The underlying problem is in core 
automatic schema merging.
   
   ### Minimal reproduce step
   
    1. Create a path-based primary-key table with:
       - id INT
       - PRIMARY KEY (id)
       - 16 buckets
    
    2. Write id = -100000.
    
    3. Append a Spark DataFrame where id has LongType, using:
       write.merge-schema=true
       write.merge-schema.type-widening=true
    
    4. Read the table and inspect its schema and bucket assignments.
   
   Automatic merging can commit  id BIGINT  while retaining  primaryKeys=[id] .
   
   The normal explicit schema-change path rejects type changes to primary-key 
and partition columns.  SchemaMergingUtils , however, merges every field type 
and preserves the existing key-name lists without validating protected fields. 
A direct filesystem schema manager can then commit the merged schema.
   
   The physical hash representation changes with the type. For example, the 
reviewed reproduction produced:
   
   ```
    INT(-100000):    bucket 3
    BIGINT(-100000): bucket 13
   ```
   The same logical key can therefore be written to two buckets and escape 
per-bucket deduplication.
   
   ### What doesn't meet your expectations?
   
   Automatic schema merging should preserve the same key-type invariants as 
explicit ALTER TABLE operations.
    
    Type changes to primary-key and partition columns should be rejected before 
a new schema is committed. Non-key widening should continue to work.
   
   ### Anything else?
   
    A shared validation method could compare the old and merged types of all 
primary-key and partition fields before commit and produce an error naming the 
protected column.
    
    Suggested tests:
    
    1. Reject widening an INT primary key to BIGINT
    2. Reject widening a partition column
    3. Verify no schema ID or schema file is created after rejection
    4. Verify non-key widening still succeeds
    5. Spark path-table integration test
    6. Regression with a key whose INT and BIGINT representations hash to 
different buckets
    
    This issue has a high correctness impact. I would reproduce it in an 
integration test once more before filing and mention whether catalog-backed 
tables follow the same path.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to