TheR1sing3un opened a new pull request, #9950: URL: https://github.com/apache/paimon/pull/9950
### Purpose Add `search_vectors(...).to_arrow(execution="ray", concurrency=..., ray_remote_args=...)` for data-evolution tables. Each Ray task handles the whole query batch for one split: index workers reuse an open shard across bounded query blocks, and raw workers stream each complete data split once with an independent top-k per query. Reuse the existing global candidate selection, driver-side batch refinement, and shared final row lookup. All queries and workers use one pinned snapshot, preserving input order, persisted index metrics, filters, deletions, and deterministic duplicate precedence. Validate NaN scores before worker or refinement top-k truncation, and reuse the existing bounded task scheduling and cancellation. Local execution remains the default. ### Tests - Python 3.11 / Ray 2.54.0: **154 passed** across the new batch Ray suite, existing single-vector Ray suite, batch raw scan, streaming refinement, and shared lookup suites. - Ray 2.44.0 / NumPy 1.24.3 / PyArrow 18.1.0: **80 passed** across both Ray vector search suites, including real workers and native vector indexes. - Cover query blocking and reader cleanup, one task per data split, one final lookup, independent global candidate sets, reversed completion order, empty and historical snapshots, concurrent commits, filters, and invalid inputs. Add raw batch coverage to the CI Ray compatibility matrix. - Flake8, Python 3.6 syntax parsing, and `git diff --check` passed. Upstream Rust planner validation is delegated to CI. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
