XiaoHongbo-Hope opened a new pull request, #10386:
URL: https://github.com/apache/paimon/pull/10386
## Purpose
Decouple Python MAP shared-shredding conversion batches from Parquet row
groups.
The writer currently calls `write_table` for every conversion batch (at most
1,024 rows), so bounded conversion unintentionally fragments the file.
## Changes
- Keep bounded conversion and accumulate physical Arrow chunks independently.
- Use `file.block-size` (128 MiB default) as an Arrow-byte buffering target
and
cap each group at 1,048,576 rows. Flush the tail at end of file.
- Split batches at the target; allow one indivisible oversized row. Preserve
shredding metadata, placement state, logical results and failure cleanup.
- Leave Native writes, file rolling, compression and page-index settings
unchanged.
The byte target is not compressed Parquet size or a process RSS limit. The
input,
current conversion batch and Parquet encoder also consume memory. Buffering
adds
up to a row-group target compared with the previous immediate-write path.
## Validation
- Python 3.11 / PyArrow 19.0.1: 102 tests and 46 subtests passed across MAP
writes/reads/projections, page-index writes and chunked row-ID updates.
- Tests cover byte/row caps, tail groups, empty input, NULLs, oversized rows,
generator cleanup, write failure and physical metadata preservation.
- Synthetic 214,016-row layout comparison: 209 row groups with the old
conversion-boundary grouping versus 1 with buffering; full logical readback
matched. This is a layout check, not an end-to-end throughput claim.
- Flake8 and `git diff --check` passed.
Draft: PyArrow 6 compatibility and representative wide-MAP peak-memory
testing
still need validation. No claim that a single row group is universally
optimal;
existing files are not rewritten by this change.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]