TheR1sing3un opened a new pull request, #9988: URL: https://github.com/apache/paimon/pull/9988
### Purpose Ray batch vector queries currently read and score refinement candidates on the driver. Run this work on Ray workers, with one task per complete data read split. The driver still selects each query's global candidates before refinement and merges the per-query top-k scores afterward. Reuse the existing batch refinement scorer to read the candidate union once per split and score each row only for the queries that selected it. All stages use the same pinned snapshot; the shared final row lookup remains on the driver. ### Tests - Batch search and new refinement tests: 48 passed on Ray 2.54 / NumPy 2.4 / Arrow 19, and 48 passed on Ray 2.44 / NumPy 1.24 / Arrow 18. - Existing single-vector search/refinement and local batch streaming refinement regressions: 67 passed. - Coverage includes all three metrics, concurrency 1/2, duplicate queries, per-query candidate membership, nullable vectors, reader cleanup on failure, empty candidates, historical snapshots/tags and an indexed-column update between candidate selection and refinement. The distributed test rejects candidate-vector reads on the driver. - Flake8, Python 3.6 syntax parsing and `git diff --check` passed for changed files. Native index tests used paimon-vindex 0.4.0. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
