TheR1sing3un opened a new pull request, #10232:
URL: https://github.com/apache/paimon/pull/10232

   ### Purpose
   
   Chunk-shuffle planning retains all temporary chunk/segment descriptors while
   constructing the output splits, increasing peak memory when a scan produces
   many chunks. Stream each group's chunks, store private descriptors in slotted
   dataclasses, and release consumed descriptors as splits are materialized.
   Data Evolution planning also releases consumed aligned-group descriptors.
   
   Preserve the existing seeded shuffle order, balanced worker assignment,
   visible row ranges, and BLOB siblings. Deletion vectors are still read once
   per relevant file or anchor per planning call. The implementation continues
   to materialize a global chunk permutation; memory remains O(number of 
chunks).
   
   ### Tests
   
   - 86 tests passed on Python 3.11.15 / PyArrow 19.0.1 / pytest 7.4.4:
     `python -m pytest pypaimon/tests/scanner 
pypaimon/tests/reader_split_generator_test.py 
pypaimon/tests/data_evolution_split_generator_test.py -q`
   - New order-compatibility cases cover three seeds, one/three/six workers,
     empty workers, null partitions, reversed manifest input, cross-file chunks,
     unknown deletion-vector cardinality, BLOB siblings, and row-ID gaps.
     Both new cases also pass against the unmodified base implementation.
   - Changed files pass flake8 using `dev/cfg.ini`; all 840 Python files pass 
the
     license-header check; `git diff --check` passes.
   - Python 3.6 grammar check passes for changed files. This is a syntax check,
     not a Python 3.6 runtime validation. Rust-native and full repository suites
     were not run.
   
   ### Benchmark
   
   Compared with base commit `a36c25306f09aabe31f88019accf0b5a1b444a8f`.
   
   Synthetic file metadata, chunk size 100, seed 42, shard index 0. Times are
   medians of five runs without tracemalloc. Peak Python allocations are 
measured
   in a separate traced run; they include returned splits and exclude imports
   and prebuilt input metadata. These numbers are not process RSS or data-read
   throughput. All eight before/after output fingerprints and split counts 
match.
   
   | Mode | Files | Rows/file | Workers | Peak Python MiB, before → after | 
Reduction | Seconds, before → after |
   |---|---:|---:|---:|---:|---:|---:|
   | append | 100 | 100000 | 1 | 103.74 → 53.41 | 48.5% | 0.611 → 0.364 |
   | append | 100 | 100000 | 8 | 51.11 → 42.82 | 16.2% | 0.305 → 0.273 |
   | append | 10000 | 100 | 1 | 9.83 → 5.52 | 43.8% | 0.059 → 0.034 |
   | append | 10000 | 100 | 8 | 5.29 → 4.46 | 15.7% | 0.025 → 0.023 |
   | data-evolution | 100 | 100000 | 1 | 80.90 → 45.04 | 44.3% | 0.497 → 0.385 |
   | data-evolution | 100 | 100000 | 8 | 51.14 → 42.84 | 16.2% | 0.328 → 0.280 |
   | data-evolution | 10000 | 100 | 1 | 9.70 → 6.12 | 37.0% | 0.061 → 0.062 |
   | data-evolution | 10000 | 100 | 8 | 7.46 → 5.89 | 21.0% | 0.057 → 0.051 |
   
   The small-file, single-worker Data Evolution timing is approximately 
unchanged;
   the intended improvement is lower planning memory, not a universal speedup.
   
   Reproduce from `paimon-python`:
   
   ```sh
   python -m pypaimon.benchmark.chunk_shuffle --mode append --files 100 
--rows-per-file 100000 --shards 1 --repeats 5
   python -m pypaimon.benchmark.chunk_shuffle --mode data-evolution --files 100 
--rows-per-file 100000 --shards 8 --repeats 5
   ```
   
   To measure the base revision, run the candidate's standalone benchmark script
   with `PYTHONPATH` pointing to a checkout of the base revision's 
`paimon-python`
   directory. Use the same interpreter and arguments. The benchmark requires no
   data files or external services.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to