TheR1sing3un opened a new pull request, #9750:
URL: https://github.com/apache/paimon/pull/9750

   ### Purpose
   
   Building a vindex index with a training sample previously loaded the entire 
temporary vector file into a NumPy array before selecting training rows. A 
small sample therefore still required memory proportional to the full index 
shard.
   
   Read selected training vectors in blocks of at most 10,000 rows instead. 
Preserve the existing sample count, evenly spaced sample positions, float32 
layout, and full-vector indexing after training. Skip gaps between read blocks 
when no training vectors are needed. The full-sample path still reads the 
complete training array required by the native API.
   
   This change only bounds the sampling-stage read buffers. It does not bound 
upstream Arrow shard materialization or the native trainer's own allocations.
   
   ### Tests
   
   - `python -m pytest pypaimon/tests/global_index_build_test.py -k vindex -q`: 
10 passed, 18 deselected. Covers sample parity, bounded reads, sparse samples, 
null-vector exclusion, full-vector indexing, and temporary-file cleanup.
   - `python -m flake8 --config dev/cfg.ini 
pypaimon/globalindex/vindex/vindex_vector_index_writer.py 
pypaimon/tests/global_index_build_test.py`: passed.
   - Native paimon-vindex 0.4.0 smoke comparison: 257 eight-dimensional 
vectors, IVF-Flat with four lists and 25% training sampling. Eight queries with 
nprobe=4 produced identical top-5 IDs and distances for old and new sampling 
paths.
   - Isolated sampling memory comparison, in separate macOS processes: a 256 
MiB file containing 262,144 256-dimensional float32 vectors, sampled at 1%, 
produced identical sample SHA-256 hashes. Peak process RSS was 279.64 MiB 
before and 33.41 MiB after. This is a sampling-only measurement, not an 
end-to-end index build benchmark; filesystem cache was not controlled.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to