TheR1sing3un opened a new pull request, #9752: URL: https://github.com/apache/paimon/pull/9752
### Purpose Vindex builds currently convert Arrow vector columns to nested Python lists, validate every float in Python, then serialize and write each row separately. This adds substantial CPU overhead even when the vectors are already stored as contiguous float32 values. Add an Arrow-aware `write_batch()` fast path for float32 List, LargeList and FixedSizeList arrays with int64 row IDs. Validate dimensions and finite values in batches and write row IDs and vector buffers through memoryviews. Handle sliced arrays and null parent vectors correctly. Invalid vectors and unsupported layouts fall back to the existing scalar path, preserving error messages and successfully written prefixes. The builder feeds bounded batches (at most 10,000 rows) into the new path, applies shard filtering and relative IDs, and rejects null row IDs before writing any batch. The scalar writer interface remains available. **This is an independent PR directly based on master (a11ce79d1). It does not include or require #9750 or #9751.** Whole-shard Arrow reading and the original training sampler remain in place; full-text and sorted-index writing are unchanged. ### Validation - `python -m pytest pypaimon/tests/vindex_batch_write_test.py pypaimon/tests/global_index_build_test.py -q`: **36 passed**. - Changed files pass `flake8 --config dev/cfg.ini` and `git diff --check`. - Byte-for-byte scalar/batch parity tests cover all three list layouts, nonzero slice offsets, null parent vectors with invalid hidden children, negative zero/subnormals, large row IDs, mixed scalar/batch calls and fallback layouts. - Invalid dimensions, null elements and NaN/Inf produce the same scalar errors and written prefixes. Tests also cover range filtering, null row-ID precedence, temporary-file cleanup and existing global-index builds. - All benchmark modes produce identical vector and row-ID file SHA-256 hashes for each dataset. - Six native builds produce identical top-10 IDs and distances for 16 fixed queries with nprobe=4 and nprobe=16. ### Benchmark and ablation Reproduction script: `paimon-python/dev/benchmark_vindex_batch_write.py`. Environment: macOS 26.4.1 arm64, Python 3.9.6, NumPy 2.0.2, PyArrow 19.0.1, paimon-vindex 0.4.0; CPU, OMP_NUM_THREADS=1, OPENBLAS_NUM_THREADS=1. Three fresh sequential processes per configuration, shuffled order, 36 runs total. Dataset preparation is separate. Filesystem cache is uncontrolled; Arrow reported sandbox restrictions on some sysctl CPU probes in all modes. Primary dataset: one index shard in a real Paimon data-evolution Parquet table, row tracking enabled, 65,536 valid non-null vectors of 256 float32 elements (64 MiB raw vectors), batch size 1,024. Ingestion-only measurements include reading, conversion/validation and temporary-file writing, but skip native training/serialization. Hash verification time is excluded. Values are medians of three runs; RSS is absolute process peak, including roughly 150 MiB of import/runtime overhead. | Mode | Conversion / writing | Time (s) | Peak RSS (MiB) | | --- | --- | ---: | ---: | | baseline | Whole-shard Python lists, scalar validation and writes | 3.300 | 941.86 | | scalar-batches | Per-batch Python lists, scalar validation and writes | 3.125 | 300.06 | | convert-only | Columnar conversion/validation, per-row file writes | 0.186 | 300.02 | | batch | Columnar conversion/validation, batch file writes | 0.123 | 299.95 | Columnar preparation accounts for most of the throughput gain; batching file writes adds another reduction. The ablation splits buffer writes into per-row **buffered Python file calls**, not OS syscalls. It compares controlled pipeline variants rather than inserting timers into per-element loops. The complete native IVF-Flat build (nlist=16, 25% training sample, L2) takes **3.419 → 0.228 s**, with peak RSS **942.23 → 507.19 MiB**. This includes training and index-file serialization, excludes hash verification, snapshot commit and subsequent query validation, and is about 15.0x faster in this local experiment. A second dataset of 16,384 × 768 float32 vectors (48 MiB) gives ingestion times of 2.347 / 2.329 / 0.123 / 0.081 s for the four modes above. Batch ingestion on the primary dataset takes 0.146 / 0.123 / 0.119 s at batch sizes 256 / 1,024 / 4,096. These results do not establish equivalent speedups for cold disk, remote storage or GPU workloads. Whole-shard Arrow buffers and native training/index allocations still contribute to memory use. Example reproduction from `paimon-python`: ```sh PYTHONPATH=. python dev/benchmark_vindex_batch_write.py prepare --warehouse /tmp/paimon-batch-bench --rows 65536 --dimension 256 OMP_NUM_THREADS=1 OPENBLAS_NUM_THREADS=1 PYTHONPATH=. python dev/benchmark_vindex_batch_write.py run --warehouse /tmp/paimon-batch-bench --mode batch --batch-size 1024 --output /tmp/batch.json ``` Run each of `baseline`, `scalar-batches`, `convert-only` and `batch` three times in fresh processes. Add `--native` to include native building and query validation. Use a separate warehouse with `--rows 16384 --dimension 768` for the high-dimensional dataset. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
