XiaoHongbo-Hope opened a new pull request, #10386:
URL: https://github.com/apache/paimon/pull/10386

   ## Purpose
   
   Decouple Python MAP shared-shredding conversion batches from Parquet row 
groups.
   The writer currently calls `write_table` for every conversion batch (at most
   1,024 rows), so bounded conversion unintentionally fragments the file.
   
   ## Changes
   
   - Keep bounded conversion and accumulate physical Arrow chunks independently.
   - Use `file.block-size` (128 MiB default) as an Arrow-byte buffering target 
and
     cap each group at 1,048,576 rows. Flush the tail at end of file.
   - Split batches at the target; allow one indivisible oversized row. Preserve
     shredding metadata, placement state, logical results and failure cleanup.
   - Leave Native writes, file rolling, compression and page-index settings 
unchanged.
   
   The byte target is not compressed Parquet size or a process RSS limit. The 
input,
   current conversion batch and Parquet encoder also consume memory. Buffering 
adds
   up to a row-group target compared with the previous immediate-write path.
   
   ## Validation
   
   - Python 3.11 / PyArrow 19.0.1: 102 tests and 46 subtests passed across MAP
     writes/reads/projections, page-index writes and chunked row-ID updates.
   - Tests cover byte/row caps, tail groups, empty input, NULLs, oversized rows,
     generator cleanup, write failure and physical metadata preservation.
   - Synthetic 214,016-row layout comparison: 209 row groups with the old
     conversion-boundary grouping versus 1 with buffering; full logical readback
     matched. This is a layout check, not an end-to-end throughput claim.
   - Flake8 and `git diff --check` passed.
   
   Draft: PyArrow 6 compatibility and representative wide-MAP peak-memory 
testing
   still need validation. No claim that a single row group is universally 
optimal;
   existing files are not rewritten by this change.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to