LuciferYang opened a new issue, #10253:
URL: https://github.com/apache/paimon/issues/10253

   ### Search before asking
   
   - [X] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   ### Paimon version
   
   master
   
   ### Compute Engine
   
   Flink / Spark (Iceberg compatibility + data evolution)
   
   ### Minimal reproduce step
   
   When a data file records `writeCols` (a partial write whose columns are a 
subset or a reorder of the table fields) and its `valueStatsCols` is null, 
`IcebergDataFileMeta.create` maps the stats row by the full Iceberg schema 
order. A null `valueStatsCols` actually means the stats cover the whole write 
schema in write-column order, which differs from the table order for a partial 
write (for example a MERGE INTO whose SET clause lists columns differently, or 
a nested write that records leaf paths).
   
   `IcebergCommitCallback` exports these files (data-evolution splits are 
always raw-convertible), so the bounds and null counts drift onto the wrong 
Iceberg field ids. For a strict-subset write the code also reads past the end 
of the stats row, producing a garbage bound or an out-of-range access.
   
   ### What doesn't meet your expectations?
   
   Iceberg readers use these stats for data skipping, so wrong 
bounds/null-counts can drop or misjudge matching rows. The stats-column list 
should be derived from the recorded write columns in their order, mapping 
nested leaf paths to their top-level field, so slot i of the stats row is 
attributed to the correct column.
   
   ### Anything else?
   
   _No response_
   
   ### Are you willing to submit a PR?
   
   - [X] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to