SteNicholas opened a new issue, #405: URL: https://github.com/apache/paimon-cpp/issues/405
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation Sub-issue of #399 (step 4: row filters). In Java, a full-text search can carry a row filter (apache/paimon#9855). This is `FullTextSearchBuilder#withFilter(Predicate)`, a Spark `WHERE` clause on `full_text_search`, or `withFilter` on a hybrid search. Only matching rows are ranked, so the result is the top k among matching rows, not a filtered subset of the unfiltered top k. The matching set is passed to the existing index as include row ids, so BM25 statistics stay those of the full corpus. Paimon C++ has no equivalent. A caller has to filter after ranking and gets fewer than `limit` rows. Java design: - **Predicate split.** `withFilter` splits the predicate. Partition predicates are ANDed into the partition filter; the rest become the data filter. - **Scan.** The scan also keeps index files that contain any filter field. - Scalar pre-filter files are files containing a filter field that are not full-text indexes, or are full-text indexes with extra fields. - Each `IndexFullTextSearchSplit` carries the scalar files that intersect its physical range. - **`DataEvolutionFullTextRead#matchedRows`** resolves the data filter into an exact set of row ids before ranking: 1. `covered` = the union of all splits' `searchRowRanges`, ANDed with the live rows. 2. `unindexed` = the coverage gaps of the scalar indexes for the filter fields, ANDed with `covered`. This is empty in `fast` mode; see {{S7}}. 3. The rows in `covered − unindexed` are evaluated with the scalar global index (`scanWithCoverage`), which returns a result plus the ids of the fields that contributed to it. 4. If that evaluation is not exact, meaning the contributing fields don't cover every filter field or there is a `Contains`/`EndsWith`/`Like` leaf, the read either: - reads the filter columns plus `_ROW_ID` of the candidate rows and verifies them (`global-index.filter.refine-from-data = true`, via `FilteredRowIdReader`), or - drops those candidates and logs a warning (the default, `false`). 5. `unindexed` rows are resolved with `FilteredRowIdReader`. - **Include set.** Each split searches with include row ids `searchRowRanges ∧ liveRows ∧ matchedRows`. ### Solution - Add a `WithFilter` equivalent to the full-text search builder, and split the predicate into partition and data filters. - Select the scalar pre-filter files in the scan and attach them to the splits. - Resolve matched rows as described above. Add the option `global-index.filter.refine-from-data` (Boolean, default `false`) with the Java description. Add a filtered row-id reader that reads only the filter columns plus `_ROW_ID`, pinned to the plan snapshot. - Extend `GlobalIndexEvaluatorImpl` to report which fields contributed to a result, so exactness can be decided. apache/paimon#8547 also fixed scalar-filter coverage for multi-field indexes. - Add tests aligned with `NativeFullTextRowFilterTest`: - equality, range and `IN` filters - a filter on a column without an index - `Like`/`Contains` with refine on and off - partition and data predicates combined ### Anything else? - Depends on {{S5}}. - Java primary-key full-text search rejects non-partition filters with `UnsupportedOperationException("Primary-key full-text search does not support non-partition filters yet.")`; see {{S11}}. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
