SteNicholas opened a new issue, #405:
URL: https://github.com/apache/paimon-cpp/issues/405

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   ### Motivation
   
   Sub-issue of #399 (step 4: row filters).
   
   In Java, a full-text search can carry a row filter (apache/paimon#9855). 
This is `FullTextSearchBuilder#withFilter(Predicate)`, a Spark `WHERE` clause 
on `full_text_search`, or `withFilter` on a hybrid search. Only matching rows 
are ranked, so the result is the top k among matching rows, not a filtered 
subset of the unfiltered top k. The matching set is passed to the existing 
index as include row ids, so BM25 statistics stay those of the full corpus.
   
   Paimon C++ has no equivalent. A caller has to filter after ranking and gets 
fewer than `limit` rows.
   
   Java design:
   
   - **Predicate split.** `withFilter` splits the predicate. Partition 
predicates are ANDed into the partition filter; the rest become the data filter.
   - **Scan.** The scan also keeps index files that contain any filter field.
     - Scalar pre-filter files are files containing a filter field that are not 
full-text indexes, or are full-text indexes with extra fields.
     - Each `IndexFullTextSearchSplit` carries the scalar files that intersect 
its physical range.
   - **`DataEvolutionFullTextRead#matchedRows`** resolves the data filter into 
an exact set of row ids before ranking:
     1. `covered` = the union of all splits' `searchRowRanges`, ANDed with the 
live rows.
     2. `unindexed` = the coverage gaps of the scalar indexes for the filter 
fields, ANDed with `covered`. This is empty in `fast` mode; see {{S7}}.
     3. The rows in `covered − unindexed` are evaluated with the scalar global 
index (`scanWithCoverage`), which returns a result plus the ids of the fields 
that contributed to it.
     4. If that evaluation is not exact, meaning the contributing fields don't 
cover every filter field or there is a `Contains`/`EndsWith`/`Like` leaf, the 
read either:
        - reads the filter columns plus `_ROW_ID` of the candidate rows and 
verifies them (`global-index.filter.refine-from-data = true`, via 
`FilteredRowIdReader`), or
        - drops those candidates and logs a warning (the default, `false`).
     5. `unindexed` rows are resolved with `FilteredRowIdReader`.
   - **Include set.** Each split searches with include row ids `searchRowRanges 
∧ liveRows ∧ matchedRows`.
   
   ### Solution
   
   - Add a `WithFilter` equivalent to the full-text search builder, and split 
the predicate into partition and data filters.
   - Select the scalar pre-filter files in the scan and attach them to the 
splits.
   - Resolve matched rows as described above. Add the option 
`global-index.filter.refine-from-data` (Boolean, default `false`) with the Java 
description. Add a filtered row-id reader that reads only the filter columns 
plus `_ROW_ID`, pinned to the plan snapshot.
   - Extend `GlobalIndexEvaluatorImpl` to report which fields contributed to a 
result, so exactness can be decided. apache/paimon#8547 also fixed 
scalar-filter coverage for multi-field indexes.
   - Add tests aligned with `NativeFullTextRowFilterTest`:
     - equality, range and `IN` filters
     - a filter on a column without an index
     - `Like`/`Contains` with refine on and off
     - partition and data predicates combined
   
   ### Anything else?
   
   - Depends on {{S5}}.
   - Java primary-key full-text search rejects non-partition filters with 
`UnsupportedOperationException("Primary-key full-text search does not support 
non-partition filters yet.")`; see {{S11}}.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to