SteNicholas opened a new issue, #410: URL: https://github.com/apache/paimon-cpp/issues/410
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation Sub-issue of #399 (step 5: primary-key full-text index, read side). Paimon C++ cannot search primary-key full-text indexes: - `PrimaryKeySortedIndexScan` rejects full-text search (`src/paimon/core/table/source/primary_key_sorted_index_scan.cpp:288-290`). - There is no scan, split, bucket search or read for `full-text` payloads. - Indexed scores are not propagated through the primary-key physical-position read path (the TODO in `src/paimon/core/operation/raw_file_split_read.cpp:85-89`). Java design (apache/paimon#8649, apache/paimon#8652, apache/paimon#8659, apache/paimon#8844, apache/paimon#9060, apache/paimon#9184): - **Dispatch.** `FullTextSearchBuilderImpl` routes to the primary-key path when the table is not a data-evolution table and `pk-full-text.index.columns` covers the column. A non-partition filter on that path is rejected with `Primary-key full-text search does not support non-partition filters yet.` - **`PrimaryKeyFullTextScan`:** 1. Plan the primary-key batch scan with the partition filter, pinned to the snapshot. 2. Scan the index manifest for `full-text` payloads that have source metadata and the definition's field id. All entries must be `ADD`. 3. Group data splits by (partition, bucket), skipping `bucket < 0`. 4. Keep eligible files with their aligned `DeletionFile`s, and resolve current payloads with `PkFullTextBucketIndexState#fromActiveDataFiles`. 5. Emit one split per bucket. - **`PrimaryKeyFullTextSearchSplit`** holds the data split, the payload files and the uncovered data file names. Every eligible file is either covered by exactly one payload or listed as uncovered. - **`PrimaryKeyFullTextBucketSearch`** (`searchRankingsAsync`): - For each payload, lay out its source files in order to get their row offsets. - If a source is inactive or has deletions, the include set is the active ranges minus deleted positions. A payload whose include set is empty is skipped. - Call `visitFullTextSearch(new FullTextSearch(column, query, limit).withIncludeRowIds(include))`. - Map hits to `PrimaryKeySearchPosition(partition, bucket, fileName, rowId - offset, score)`, sorted by score descending, then file name, then position. - **`PrimaryKeyFullTextRead`:** - Requires `limit > 0`, and supports only `full-text-index.search-mode = fast`; `full`/`detail` throw `UnsupportedOperationException`. - Searches each split asynchronously on the `global-index.thread-num` executor, submitting from the caller thread. - Takes the global top-k with `PrimaryKeySearchRanker#topKByScore`. - Returns `PrimaryKeyScoredResult`. It turns into one indexed split per data file, with row ranges, scores and that file's `DeletionFile`, which are read by position. - Uncovered data files are not searched in `fast` mode. ### Solution - Port `PrimaryKeyFullTextScan`, `PrimaryKeyFullTextSearchSplit` (serializable), `PrimaryKeyFullTextBucketSearch` and `PrimaryKeyFullTextRead`. - Port the shared primitives `PrimaryKeySearchPosition`, `PrimaryKeySearchRanker#topKByScore` and `PrimaryKeyScoredResult`. Hybrid search ({{S8}}) will reuse them. - Propagate `_INDEX_SCORE` through the primary-key positional read path. - Add primary-key dispatch to the table-level full-text builder from {{S5}}. - Add tests aligned with Java `PrimaryKeyFullTextScanTest`, `PrimaryKeyFullTextReadTest`, `PrimaryKeyFullTextSearchTest`, `PrimaryKeyFullTextBucketSearchTest` and `NativePrimaryKeyFullTextIndexTest`: - deletion vectors - several buckets and levels - uncovered files - global top-k - partition filters - rejection of non-partition filters and non-`fast` modes - archives written by Java ### Anything else? Depends on {{S5}} and {{S10}}. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
