TheR1sing3un opened a new pull request, #9951: URL: https://github.com/apache/paimon/pull/9951
### Purpose Single-vector Ray search currently reads and scores refinement candidates on the driver. Distribute this phase through the existing raw-scan workers: after selecting the global approximate top-(k × refine_factor), intersect its row-ID ranges with the planned read ranges and dispatch complete data splits. Each worker streams and scores only those candidates, returning local top-k row IDs and scores for the final merge. This preserves global candidate membership, the persisted index metric, deterministic ranking, filters, deletions, and the pinned read snapshot while moving refinement vector reads and scoring off the driver. Final row lookup stays on the driver. Reuse the existing task concurrency, retry, cancellation, and reader cleanup paths. ### Tests - Python 3.11 / Ray 2.54.0: **50 passed** across the new refinement and existing Ray vector search suites, including native vector indexes and real workers. - Ray 2.44.0 / NumPy 1.24.3 / PyArrow 18.1.0: **11 passed** in the new refinement suite. - Verify all three metrics and concurrency 1/2 across real splits while rejecting vector reads on the driver; global candidate selection under reversed completion order; exact candidate intersection and NaN rejection; empty candidates; historical snapshots and retained tags with filters/deletions; and an indexed-column update between candidate selection and refinement. - Flake8, Python 3.6 syntax parsing, and `git diff --check` passed. Upstream Rust planner validation is delegated to CI. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
