TheR1sing3un opened a new pull request, #10232:
URL: https://github.com/apache/paimon/pull/10232
### Purpose
Chunk-shuffle planning retains all temporary chunk/segment descriptors while
constructing the output splits, increasing peak memory when a scan produces
many chunks. Stream each group's chunks, store private descriptors in slotted
dataclasses, and release consumed descriptors as splits are materialized.
Data Evolution planning also releases consumed aligned-group descriptors.
Preserve the existing seeded shuffle order, balanced worker assignment,
visible row ranges, and BLOB siblings. Deletion vectors are still read once
per relevant file or anchor per planning call. The implementation continues
to materialize a global chunk permutation; memory remains O(number of
chunks).
### Tests
- 86 tests passed on Python 3.11.15 / PyArrow 19.0.1 / pytest 7.4.4:
`python -m pytest pypaimon/tests/scanner
pypaimon/tests/reader_split_generator_test.py
pypaimon/tests/data_evolution_split_generator_test.py -q`
- New order-compatibility cases cover three seeds, one/three/six workers,
empty workers, null partitions, reversed manifest input, cross-file chunks,
unknown deletion-vector cardinality, BLOB siblings, and row-ID gaps.
Both new cases also pass against the unmodified base implementation.
- Changed files pass flake8 using `dev/cfg.ini`; all 840 Python files pass
the
license-header check; `git diff --check` passes.
- Python 3.6 grammar check passes for changed files. This is a syntax check,
not a Python 3.6 runtime validation. Rust-native and full repository suites
were not run.
### Benchmark
Compared with base commit `a36c25306f09aabe31f88019accf0b5a1b444a8f`.
Synthetic file metadata, chunk size 100, seed 42, shard index 0. Times are
medians of five runs without tracemalloc. Peak Python allocations are
measured
in a separate traced run; they include returned splits and exclude imports
and prebuilt input metadata. These numbers are not process RSS or data-read
throughput. All eight before/after output fingerprints and split counts
match.
| Mode | Files | Rows/file | Workers | Peak Python MiB, before → after |
Reduction | Seconds, before → after |
|---|---:|---:|---:|---:|---:|---:|
| append | 100 | 100000 | 1 | 103.74 → 53.41 | 48.5% | 0.611 → 0.364 |
| append | 100 | 100000 | 8 | 51.11 → 42.82 | 16.2% | 0.305 → 0.273 |
| append | 10000 | 100 | 1 | 9.83 → 5.52 | 43.8% | 0.059 → 0.034 |
| append | 10000 | 100 | 8 | 5.29 → 4.46 | 15.7% | 0.025 → 0.023 |
| data-evolution | 100 | 100000 | 1 | 80.90 → 45.04 | 44.3% | 0.497 → 0.385 |
| data-evolution | 100 | 100000 | 8 | 51.14 → 42.84 | 16.2% | 0.328 → 0.280 |
| data-evolution | 10000 | 100 | 1 | 9.70 → 6.12 | 37.0% | 0.061 → 0.062 |
| data-evolution | 10000 | 100 | 8 | 7.46 → 5.89 | 21.0% | 0.057 → 0.051 |
The small-file, single-worker Data Evolution timing is approximately
unchanged;
the intended improvement is lower planning memory, not a universal speedup.
Reproduce from `paimon-python`:
```sh
python -m pypaimon.benchmark.chunk_shuffle --mode append --files 100
--rows-per-file 100000 --shards 1 --repeats 5
python -m pypaimon.benchmark.chunk_shuffle --mode data-evolution --files 100
--rows-per-file 100000 --shards 8 --repeats 5
```
To measure the base revision, run the candidate's standalone benchmark script
with `PYTHONPATH` pointing to a checkout of the base revision's
`paimon-python`
directory. Use the same interpreter and arguments. The benchmark requires no
data files or external services.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]