zhuxiangyi opened a new pull request, #9953: URL: https://github.com/apache/paimon/pull/9953
### Purpose Vector search with `withFilter` hands the scalar global index answer to the ANN as the include bitmap (`AbstractDataEvolutionVectorRead.scalarMatchedRows`). Global index answers are only candidates, not exact matches: BTree answers `contains` / `endsWith` / `like` with every non-null row, and `GlobalIndexEvaluator` drops a conjunct no index can evaluate. The ANN then ranks that superset, a non-matching but closer row takes a top-k slot, and the engine-side filter afterwards cannot bring the dropped matching row back — the user sees fewer rows than exist, possibly none. Reproduced on master with `(id INT, name STRING, vec ARRAY<FLOAT>)`, rows `alpha (1,0)`, `beta zeta (0.6,0.8)`, `gamma (0,1)`, query `(1,0)`, `limit 1`: - `withFilter(contains(name, 'zeta'))` with a BTree on `name` returns row 0 instead of row 1. - `withFilter(id >= 0 AND name = 'beta zeta')` with a BTree on `id` only returns row 0 instead of row 1. This is the mechanism the review of #9855 caught in the new full-text pre-filter; the vector pre-filter has had it since #8459. #9909 handled the case where no index can evaluate the filter at all; this PR handles the case where an index answers with a superset. ### Changes - `AbstractDataEvolutionVectorRead.scalarMatchedRows` uses `scanWithCoverage` and, when the answer is not exact, refines the candidates through `FilteredRowIdReader` (reads only the filter columns of the candidate rows, bounded to the split ranges, and evaluates the predicate on the data). Exact answers are used as before. Batch vector search and hybrid vector routes share this path. - The exactness check (`contributingFieldIds` covers every predicate field, no `Contains` / `EndsWith` / `Like` leaf) moves from `DataEvolutionFullTextRead` to `FilteredRowIdReader.isExact`, so full-text and vector search follow one rule. - `rawPreFilter` is unchanged: it only bounds the raw scan, and the raw read still evaluates the predicate with `executeFilter()`. ### Tests `VectorSearchRowFilterExactnessTest` (8): - `contains` on a BTree column and a partially indexed conjunction return the matching row, not the nearest neighbour - `endsWith` / `like`; a candidate set that refines to nothing; an OR with an unevaluable branch - candidates refined per index range with two vector index files - refinement combined with deletion vectors - batch vector search and a hybrid vector route - a filter column with no index at all in `fast` and `full` mode (unchanged behaviour after #9909, kept as a guard) Regression: `paimon-core` `table.source.*` + `globalindex.**` (369), `VectorSearchBuilderTest` (41), `FullTextSearchBuilderTest` (44), `PrimaryKeyVectorSearchTest` (8); Spark `VectorSearchOptionsTest`, `PrimaryKeyVectorSearchTest`, `HybridSearchTest`, `FullTextSearchTest`, `TableValuedFunctionsTest` (56). ### API and Format No API, option, or format changes. Behaviour change: a vector search whose filter can only be answered as candidates by the scalar index now reads the filter columns of those candidates before ranking, instead of ranking the superset. ### Documentation None needed; the vector search docs already describe the filter as "evaluated with matching scalar global indexes before vector search", which is now exact. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
