joseph-isaacs opened a new pull request, #10409:
URL: https://github.com/apache/paimon/pull/10409

   ### Purpose
   
   PyPaimon Vortex scans spend avoidable time converting and slicing batches, 
and Vortex writes always use the default compression preset even when users 
prioritize smaller files.
   
   Use native Arrow conversion and layout-aware Vortex scan splits, slice 
output batches without copying, and reuse Arrow batches that already match the 
requested schema. Preserve projection, filtering, indexed reads, and schema 
evolution. Upgrade vortex-data to 0.87.0.
   
   Expose `vortex.compact.enabled` (default `false`) for Python Vortex data and 
vector writes through local, PyArrow, and HDFS-native FileIO:
   
   ```python
   table = table.copy({"vortex.compact.enabled": "true"})
   ```
   
   Compact mode adaptively enables denser encodings, including Zstd for strings 
and Pco for numeric data. Storage options and failure cleanup are preserved. 
Document that `file.compression` and `file.compression.zstd-level` do not 
configure Vortex's presets. Set the ClickBench Parquet comparison to Zstd level 
3 explicitly.
   
   For the same 3 million ClickBench rows and nine file boundaries, local 
output sizes were 420.04 MiB for default Vortex, 330.01 MiB for Parquet 
Zstd(3), and 279.37 MiB for Vortex compact. The compact size measurement used 
the Vortex compact API directly; the added table-option integration is covered 
by the tests below. Compact read performance was not benchmarked.
   
   ### Tests
   
   - 106 passed: `vortex_writer_options_test.py`, `vortex_reader_test.py`, and 
`hdfs_native_test.py`.
   - Coverage includes option omitted/false/true; append, primary-key, 
data-evolution, and vector round trips; storage-option forwarding; and 
failed-write cleanup.
   - Flake8 passed on all Python files changed by the compact option, using 
`paimon-python/dev/cfg.ini`.
   - `git diff --check` passed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to