unikdahal opened a new issue, #3287:
URL: https://github.com/apache/iceberg-rust/issues/3287

   ### Is your feature request related to a problem or challenge?
   
   Iceberg Rust has basic position-delete file writing support, but row-level 
rewrite workflows also need to:
   
   - load positions from an existing file-scoped position-delete file
   - validate that the delete file actually targets the expected data file
   - merge existing and newly produced delete positions
   - write the resulting positions back in Iceberg-required "(file_path, pos)" 
order
   - de-duplicate repeated positions
   
   One motivating use case is native Iceberg merge-on-read row-level writes in 
Apache DataFusion Comet, where an execution engine may need to read existing 
position deletes for a data file, combine them with newly generated deletes, 
and write a canonical replacement delete file.
   The functionality is generic to Iceberg Rust and should also be useful for 
other row-level write and delete-file rewrite implementations.
   
   
   
   ### Describe the solution you'd like
   
   Position delete index loader
   
   Add a loader for a single file-scoped V2 Parquet position-delete file that:
   
   - resolves "file_path" and "pos" using the reserved Iceberg field IDs
   - validates the physical schema and delete-file metadata
   - verifies all physical rows target the expected data file
   - validates non-negative positions and record counts
   - canonicalizes duplicate/out-of-order positions into a sorted position index
   
   This should stay separate from the existing scan-oriented delete loader, 
since rewrite callers require the stronger invariant that the entire supplied 
delete file belongs to one data file.
   
   Sorting position-only delete writer
   
   Add a writer for unordered "(file_path, pos)" input that:
   
   - buffers positions by data-file path
   - de-duplicates positions
   - writes rows sorted by "file_path", then "pos"
   - supports both "RecordBatch" input and individual position insertion
   - preserves Iceberg position-delete metric semantics:
     - reserved-column counts are omitted
     - bounds are retained for file-scoped delete files
     - bounds are removed when a physical delete file references multiple data 
files
   
   The initial implementation can remain in-memory; spilling can be added 
separately if needed.
   
   Scope
   
   This is intended for V2 position-delete files. It does not add V3 
deletion-vector writing or transaction-level row-delta commit handling.
   
   The writer also does not need to own discovery of previous delete files. 
Callers can load existing file-scoped deletes using the position-delete index 
loader and feed those positions into the sorting writer together with newly 
generated deletes.
   
   ### Willingness to contribute
   
   I can contribute to this feature independently


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to