SteNicholas opened a new issue, #406: URL: https://github.com/apache/paimon-cpp/issues/406
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation Sub-issue of #399 (step 4: search modes). In Paimon C++, a full-text search only sees rows covered by a full-text index. Rows appended after the last index build are invisible. Java controls this with `full-text-index.search-mode`, which falls back to `global-index.search-mode`. In `full` and `detail` modes, Java also searches row ranges that have no full-text index yet: it reads the raw rows and builds a temporary index (apache/paimon#8316, apache/paimon#8844). Java options (`CoreOptions`): | Key | Default | Description | |---|---|---| | `global-index.search-mode` | (none) | Fallback search mode for global index queries. | | `scalar-index.search-mode` | `fast` (apache/paimon#8891) | Search mode for scalar index queries. | | `vector-index.search-mode` | `fast` | Search mode for vector index queries. | | `full-text-index.search-mode` | `fast` | Search mode for full-text index queries. | Values: - `fast`: search indexed data only. - `full`: use the snapshot's next row id and the global index coverage to find missing row ids, and scan raw data only when a gap exists. - `detail`: scan data files to find the exact unindexed rows. Resolution order: an explicitly set family key wins, then `global-index.search-mode`, then the family default. Java design: - **Coverage.** `DataEvolutionGlobalIndexCoverage#unindexedRanges(fieldIds, ...)` computes the gaps: - `fast`: none. The same applies when the snapshot's `nextRowId` is null or not positive. - `full`: `[0, nextRowId - 1]` minus the indexed ranges. - `detail`: the non-null row-id ranges of all data files (a `ScanMode.ALL` read that respects the partition filter), minus the indexed ranges. - Indexed ranges are intersected across the requested fields. Both `index_field_id` and `extra_field_ids` count as coverage. - **Scan.** When at least one full-text index file exists and the unindexed ranges are non-empty, the scan adds a `RawFullTextSearchSplit(rowRanges)`. With no full-text index at all, the result is empty even in `full` mode. - **Read.** `RawFullTextReadImpl` handles the raw split: 1. Read the column plus `_ROW_ID` for the raw ranges, pinned to the plan snapshot. The read respects deletion vectors, and the row filter is applied to build the include set. 2. Build a temporary in-memory index with the column's index type and options, writing `(text, rowId - first.from)`. 3. Search it with the same query and limit. 4. Replace indexed hits that fall inside the raw ranges with the raw hits. 5. Apply the final top-k. The temporary index's statistics come from the raw rows only. ### Solution - Add the four options and the family-specific resolution. - Port the coverage computation, `RawFullTextSearchSplit`, and the raw read path with the temporary index. The temporary index should be written to and read from memory, without touching table storage. - Use `scalar-index.search-mode` for row-filter coverage in {{S6}}. - Add tests: - rows appended after the index build, in `fast`, `full` and `detail` modes - partial index coverage across partitions - deletion vectors inside raw ranges - a row filter on the raw path ### Anything else? - Depends on {{S5}}. The raw path supports row filters once {{S6}} lands. - Java primary-key full-text search supports only `fast` and rejects the other modes; see {{S11}}. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
