TheR1sing3un opened a new pull request, #9988:
URL: https://github.com/apache/paimon/pull/9988

   ### Purpose
   
   Ray batch vector queries currently read and score refinement candidates on 
the driver. Run this work on Ray workers, with one task per complete data read 
split. The driver still selects each query's global candidates before 
refinement and merges the per-query top-k scores afterward.
   
   Reuse the existing batch refinement scorer to read the candidate union once 
per split and score each row only for the queries that selected it. All stages 
use the same pinned snapshot; the shared final row lookup remains on the driver.
   
   ### Tests
   
   - Batch search and new refinement tests: 48 passed on Ray 2.54 / NumPy 2.4 / 
Arrow 19, and 48 passed on Ray 2.44 / NumPy 1.24 / Arrow 18.
   - Existing single-vector search/refinement and local batch streaming 
refinement regressions: 67 passed.
   - Coverage includes all three metrics, concurrency 1/2, duplicate queries, 
per-query candidate membership, nullable vectors, reader cleanup on failure, 
empty candidates, historical snapshots/tags and an indexed-column update 
between candidate selection and refinement. The distributed test rejects 
candidate-vector reads on the driver.
   - Flake8, Python 3.6 syntax parsing and `git diff --check` passed for 
changed files. Native index tests used paimon-vindex 0.4.0.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to