LuciferYang opened a new pull request, #10255:
URL: https://github.com/apache/paimon/pull/10255

   ### Purpose
   
   `IcebergDataFileMeta.create` treated a null `statsColumns` as covering the 
whole Iceberg (table) schema in table order. A null `valueStatsCols` instead 
means the stats cover the whole write schema, whose column order and set can 
differ from the table for a partial write. As a result the min/max bounds and 
null counts drifted onto the wrong Iceberg fields in the manifest, and a 
strict-subset write could read past the end of the stats row.
   
   This passes the file's `writeCols` into `create`. When `statsColumns` is 
null and `writeCols` is present, it builds the explicit stats-column list from 
the write columns in their recorded order, mapping a nested leaf path to its 
top-level field so each stats slot is attributed to the correct column. Full 
writes keep `writeCols` null and are unaffected.
   
   This closes #10253.
   
   ### Tests
   
   
`IcebergDataFileMetaTest.testNullStatsColumnsWithWriteColsAlignByWriteSchema` 
pins that a partial write of `(b, k)` in write-column order maps the bounds and 
null counts to `b` and `k` and does not drift onto the unwritten column `a`.
   
   
`IcebergDataFileMetaTest.testNullStatsColumnsWithNestedWriteColsMapToTopLevel` 
pins that a nested leaf path attributes its stats to the top-level field rather 
than to a neighboring column.
   
   ### API and Format
   
   No.
   
   ### Documentation
   
   No.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to