TheR1sing3un opened a new pull request, #9930:
URL: https://github.com/apache/paimon/pull/9930

   ### Purpose
   
   Vector queries currently execute index-shard searches and raw-vector scans 
in the driver process. Add opt-in Ray execution for single-vector queries on 
data-evolution tables:
   
   ```python
   neighbors = (
       docs.search(query_vector, column="embedding", pre_filter="category = 
'lake'")
       .select(["id", "content"])
       .limit(10)
       .to_arrow(execution="ray", concurrency=4, ray_remote_args={"num_cpus": 
1})
   )
   ```
   
   The driver pins the read snapshot and plans filters once. Ray tasks search 
index shards or stream complete raw-data read splits, returning candidate IDs 
and scores. The existing reader retains global candidate selection, duplicate 
precedence, refinement, and final result lookup. Workers return the persisted 
index metric, including when a shard has no hits, so raw fallback uses the same 
scoring convention. Refinement and lookup remain on the driver.
   
   Task submission is bounded, and a failure cancels outstanding tasks without 
returning partial results. Native-reader initialization also releases resources 
when task cancellation interrupts it. Non-finite query vectors and NaN scores 
are rejected for Ray execution because NaN cannot participate in a consistent 
distributed top-k ordering.
   
   Local execution remains the default and does not require Ray. Primary-key, 
batch-vector, and hybrid queries are outside this change. The documentation 
covers dependencies, snapshot semantics, shared storage, and scheduling costs.
   
   ### Tests
   
   - 39 Ray query/execution tests pass locally on Ray 2.44.0 and 2.58.0, 
including real worker processes, multiple raw/index splits, all three metrics, 
global refinement, duplicate precedence, filters/deletions, historical and 
retained-tag snapshots, concurrent writes, retries, cancellation, and invalid 
inputs.
   - Existing search/multimodal regression run: 334 passed, 102 subtests 
passed. Additional reader/metric cleanup checks: 23 passed, 20 subtests passed.
   - Added raw-search smoke coverage to the existing Ray-version CI matrix. 
Additional local version checks and community CI are in progress.
   - Flake8, diff checks, and Python 3.6 syntax parsing passed.
   
   Small single-host measurement: 100,000 32-dimensional vectors, four read 
splits/index shards, K=10, one warm-up and the median of three queries. Raw 
scan was 50.19 ms locally and 26.11 ms with four Ray tasks. Fully probed 
IVF-flat was 7.13 ms locally and 9.81 ms with four Ray tasks; scheduling 
overhead dominates that small indexed query. Recall@10 was 1.0 in all cases. 
These measurements do not establish multi-node or remote-storage performance.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to