zhuxiangyi commented on PR #9953: URL: https://github.com/apache/paimon/pull/9953#issuecomment-5732507049
@JingsongLi Agreed, the refinement is a single-threaded read on the caller side, and a BTree answering `contains` makes it a whole column. I'll make it opt-in: - New option `global-index.filter.refine-from-data` (boolean, default `false`), read by vector, hybrid and full-text search. - When `false`, a scalar index answer that may be a superset (a dropped conjunct, or `contains` / `endsWith` / `like` on a BTree) is excluded instead of ranked, with a WARN. The result can be short but never contains a non-matching row; using the superset directly would return wrong rows through the Java/Python API, which has no engine-side post-filter. - When `true`, the current behaviour: refine the candidates with `FilteredRowIdReader`. - The full-text refinement path from #9855 is switched to the same option. One thing I would leave outside the option: in `scalar-index.search-mode=full`, rows whose filter columns have no index are already read from data (that replaced the temporary-index rebuild in #9855, which was more expensive). Please say if you want that gated as well. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
