TheR1sing3un opened a new pull request, #9950:
URL: https://github.com/apache/paimon/pull/9950

   ### Purpose
   
   Add `search_vectors(...).to_arrow(execution="ray", concurrency=..., 
ray_remote_args=...)` for data-evolution tables. Each Ray task handles the 
whole query batch for one split: index workers reuse an open shard across 
bounded query blocks, and raw workers stream each complete data split once with 
an independent top-k per query.
   
   Reuse the existing global candidate selection, driver-side batch refinement, 
and shared final row lookup. All queries and workers use one pinned snapshot, 
preserving input order, persisted index metrics, filters, deletions, and 
deterministic duplicate precedence. Validate NaN scores before worker or 
refinement top-k truncation, and reuse the existing bounded task scheduling and 
cancellation. Local execution remains the default.
   
   ### Tests
   
   - Python 3.11 / Ray 2.54.0: **154 passed** across the new batch Ray suite, 
existing single-vector Ray suite, batch raw scan, streaming refinement, and 
shared lookup suites.
   - Ray 2.44.0 / NumPy 1.24.3 / PyArrow 18.1.0: **80 passed** across both Ray 
vector search suites, including real workers and native vector indexes.
   - Cover query blocking and reader cleanup, one task per data split, one 
final lookup, independent global candidate sets, reversed completion order, 
empty and historical snapshots, concurrent commits, filters, and invalid 
inputs. Add raw batch coverage to the CI Ray compatibility matrix.
   - Flake8, Python 3.6 syntax parsing, and `git diff --check` passed. Upstream 
Rust planner validation is delegated to CI.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to