zhuxiangyi opened a new pull request, #9953:
URL: https://github.com/apache/paimon/pull/9953

   ### Purpose
   
   Vector search with `withFilter` hands the scalar global index answer to the 
ANN as the include bitmap 
(`AbstractDataEvolutionVectorRead.scalarMatchedRows`). Global index answers are 
only candidates, not exact matches: BTree answers `contains` / `endsWith` / 
`like` with every non-null row, and `GlobalIndexEvaluator` drops a conjunct no 
index can evaluate. The ANN then ranks that superset, a non-matching but closer 
row takes a top-k slot, and the engine-side filter afterwards cannot bring the 
dropped matching row back — the user sees fewer rows than exist, possibly none.
   
   Reproduced on master with `(id INT, name STRING, vec ARRAY<FLOAT>)`, rows 
`alpha (1,0)`, `beta zeta (0.6,0.8)`, `gamma (0,1)`, query `(1,0)`, `limit 1`:
   
   - `withFilter(contains(name, 'zeta'))` with a BTree on `name` returns row 0 
instead of row 1.
   - `withFilter(id >= 0 AND name = 'beta zeta')` with a BTree on `id` only 
returns row 0 instead of row 1.
   
   This is the mechanism the review of #9855 caught in the new full-text 
pre-filter; the vector pre-filter has had it since #8459. #9909 handled the 
case where no index can evaluate the filter at all; this PR handles the case 
where an index answers with a superset.
   
   ### Changes
   
   - `AbstractDataEvolutionVectorRead.scalarMatchedRows` uses 
`scanWithCoverage` and, when the answer is not exact, refines the candidates 
through `FilteredRowIdReader` (reads only the filter columns of the candidate 
rows, bounded to the split ranges, and evaluates the predicate on the data). 
Exact answers are used as before. Batch vector search and hybrid vector routes 
share this path.
   - The exactness check (`contributingFieldIds` covers every predicate field, 
no `Contains` / `EndsWith` / `Like` leaf) moves from 
`DataEvolutionFullTextRead` to `FilteredRowIdReader.isExact`, so full-text and 
vector search follow one rule.
   - `rawPreFilter` is unchanged: it only bounds the raw scan, and the raw read 
still evaluates the predicate with `executeFilter()`.
   
   ### Tests
   
   `VectorSearchRowFilterExactnessTest` (8):
   
   - `contains` on a BTree column and a partially indexed conjunction return 
the matching row, not the nearest neighbour
   - `endsWith` / `like`; a candidate set that refines to nothing; an OR with 
an unevaluable branch
   - candidates refined per index range with two vector index files
   - refinement combined with deletion vectors
   - batch vector search and a hybrid vector route
   - a filter column with no index at all in `fast` and `full` mode (unchanged 
behaviour after #9909, kept as a guard)
   
   Regression: `paimon-core` `table.source.*` + `globalindex.**` (369), 
`VectorSearchBuilderTest` (41), `FullTextSearchBuilderTest` (44), 
`PrimaryKeyVectorSearchTest` (8); Spark `VectorSearchOptionsTest`, 
`PrimaryKeyVectorSearchTest`, `HybridSearchTest`, `FullTextSearchTest`, 
`TableValuedFunctionsTest` (56).
   
   ### API and Format
   
   No API, option, or format changes. Behaviour change: a vector search whose 
filter can only be answered as candidates by the scalar index now reads the 
filter columns of those candidates before ranking, instead of ranking the 
superset.
   
   ### Documentation
   
   None needed; the vector search docs already describe the filter as 
"evaluated with matching scalar global indexes before vector search", which is 
now exact.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to