TheR1sing3un opened a new pull request, #9989: URL: https://github.com/apache/paimon/pull/9989
### Purpose Single and batch vector index searches currently retain every shard result before building the merged candidate maps. Local execution also submits all shard searches at once. Merge results incrementally and bound the work window in both local and Ray execution. Consume index results in plan order so the first planned occurrence of a duplicate row ID still wins. Keep global top-k and refinement candidate selection unchanged. Ray's existing completion-order scheduling remains the default for raw scans and refinement; index tasks use ordered output with pending tasks and buffered results sharing the concurrency limit. A slow earlier shard can therefore delay further submissions. The final deduplicated candidate maps still scale with unique candidate rows. ### Tests - Local and Ray vector search, batch search, refinement, parallel index search and shared batch lookup regressions: 183 passed on Ray 2.54 / NumPy 2.4 / Arrow 19. - Ray 2.44 / NumPy 1.24 / Arrow 18 compatibility: 82 single/batch Ray search tests passed. - Tests cover reverse completion order with a slow first task, bounded submissions and result retention, prompt release after merging, duplicate precedence, score ties, per-query candidates, metric mismatch, retries and failure/early-close cleanup. - Flake8, Python 3.6 syntax parsing and `git diff --check` passed for changed files. Native index tests used paimon-vindex 0.4.0. - Synthetic local batch comparison against `3ed90bc68`: 256 disjoint shards, 8 queries, 256 candidates per shard/query, concurrency 4. `tracemalloc` peak Python allocation decreased from 81.00 MiB to 46.54 MiB (42.5%), with identical result IDs and scores. This measures Python allocations, not RSS or total Ray cluster memory. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
