TheR1sing3un opened a new pull request, #9930:
URL: https://github.com/apache/paimon/pull/9930
### Purpose
Vector queries currently execute index-shard searches and raw-vector scans
in the driver process. Add opt-in Ray execution for single-vector queries on
data-evolution tables:
```python
neighbors = (
docs.search(query_vector, column="embedding", pre_filter="category =
'lake'")
.select(["id", "content"])
.limit(10)
.to_arrow(execution="ray", concurrency=4, ray_remote_args={"num_cpus":
1})
)
```
The driver pins the read snapshot and plans filters once. Ray tasks search
index shards or stream complete raw-data read splits, returning candidate IDs
and scores. The existing reader retains global candidate selection, duplicate
precedence, refinement, and final result lookup. Workers return the persisted
index metric, including when a shard has no hits, so raw fallback uses the same
scoring convention. Refinement and lookup remain on the driver.
Task submission is bounded, and a failure cancels outstanding tasks without
returning partial results. Native-reader initialization also releases resources
when task cancellation interrupts it. Non-finite query vectors and NaN scores
are rejected for Ray execution because NaN cannot participate in a consistent
distributed top-k ordering.
Local execution remains the default and does not require Ray. Primary-key,
batch-vector, and hybrid queries are outside this change. The documentation
covers dependencies, snapshot semantics, shared storage, and scheduling costs.
### Tests
- 39 Ray query/execution tests pass locally on Ray 2.44.0 and 2.58.0,
including real worker processes, multiple raw/index splits, all three metrics,
global refinement, duplicate precedence, filters/deletions, historical and
retained-tag snapshots, concurrent writes, retries, cancellation, and invalid
inputs.
- Existing search/multimodal regression run: 334 passed, 102 subtests
passed. Additional reader/metric cleanup checks: 23 passed, 20 subtests passed.
- Added raw-search smoke coverage to the existing Ray-version CI matrix.
Additional local version checks and community CI are in progress.
- Flake8, diff checks, and Python 3.6 syntax parsing passed.
Small single-host measurement: 100,000 32-dimensional vectors, four read
splits/index shards, K=10, one warm-up and the median of three queries. Raw
scan was 50.19 ms locally and 26.11 ms with four Ray tasks. Fully probed
IVF-flat was 7.13 ms locally and 9.81 ms with four Ray tasks; scheduling
overhead dominates that small indexed query. Recall@10 was 1.0 in all cases.
These measurements do not establish multi-node or remote-storage performance.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]