SteNicholas opened a new issue, #404: URL: https://github.com/apache/paimon-cpp/issues/404
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation Sub-issue of #399 (step 4: table-level full-text search). Paimon C++ only exposes the reader primitive. To run a full-text search, a caller has to: 1. call `GlobalIndexScan::CreateReaders`; 2. call `VisitFullTextSearch` on the readers; 3. pass the result to `ScanContextBuilder::SetGlobalIndexResult`. Along the way the caller also has to choose index types and shard ranges, and there is no serializable plan for running scan and read in different processes. Java offers a table-level API instead: ```java GlobalIndexResult result = table.newFullTextSearchBuilder() .withQuery("content", "{\"match\":{\"query\":\"paimon lake\"}}") .withLimit(10) .executeLocal(); TableScan.Plan plan = table.newReadBuilder().newScan().withGlobalIndexResult(result).plan(); ``` The Java components live in paimon-core `org.apache.paimon.table.source`, plus one method in paimon-common: | Java class / method | What it provides | |---|---| | `FullTextSearchBuilder` / `FullTextSearchBuilderImpl` | `withPartitionFilter`, `withFilter` (see {{S6}}), `withLimit`, `withQuery(fieldName, query)`, `newFullTextScan()`, `newFullTextRead()`, `executeLocal()`. A package-private `withSnapshot` is used by hybrid search. The query column must exist, and `newFullTextRead` requires `limit > 0`. | | `FullTextScan#scan()` | Returns `Plan { splits(), snapshot() }`; implemented by `DataEvolutionFullTextScan`. | | `FullTextRead#read(plan \| splits)` | Returns a `GlobalIndexResult`; implemented by `DataEvolutionFullTextRead`. | | `FullTextSearchSplit`, `IndexFullTextSearchSplit` | The split carries the column, the physical `rowRangeStart`/`rowRangeEnd`, the assigned `searchRowRanges`, the full-text index files and the scalar index files, with versioned serialization. | | `GlobalIndexerFactory#supportsFullTextSearch()` (paimon-common) | Defaults to `false` and is `true` for `full-text` (apache/paimon#8308). The scan uses it to decide which index files can serve full-text search. | ### Solution 1. **API.** Add C++ equivalents of the builder, scan, read and splits, following the existing builder conventions (`ScanContextBuilder` / `ReadContextBuilder`). Splits must be serializable so scan and read can run in different processes. The read result plugs into `ScanContextBuilder::SetGlobalIndexResult`, and scores stay readable through `_INDEX_SCORE`. 2. **Capability.** Add `SupportsFullTextSearch()` to `GlobalIndexerFactory` or `GlobalIndexer`, returning `true` for `full-text`. Decide whether `lucene-fts` returns `true` as well. 3. **Scan**, following `DataEvolutionFullTextScan`: - Pin the snapshot: the latest one, or the time-travel target. - Keep index manifest entries that have `GlobalIndexMeta` and pass the partition filter. - Treat a file as a full-text input when the column is its `index_field_id` or one of its `extra_field_ids`, and its factory supports full-text search. - Select ranges across overlapping and multi-field indexes (apache/paimon#8547): 1. Group candidates by column → physical range → index identity, where the identity is `type|index_field_id|extra_field_ids`. 2. Sweep the elementary intervals and give each one to the best active candidate. A dedicated index comes first, then order by identity, `from` and `to`. 3. Merge adjacent intervals. 4. Emit one `IndexFullTextSearchSplit` per selection. It keeps the physical range, which is needed for the row-id offset, and the assigned `searchRowRanges`. 4. **Read**, following `DataEvolutionFullTextRead`: - Run on an executor sized by `global-index.thread-num`. - For each split: - Create the indexer from the split's own file meta, including any extra fields, and wrap its reader in `OffsetGlobalIndexReader(start, end)`. - Search with `FullTextSearch(column, query, limit)`, restricted to `searchRowRanges ∧ liveRows`. Pass no restriction when this covers the whole physical range, and skip the split when it is empty. - Close the reader afterwards. - Merge the per-split results, keeping the first score for a duplicate row id. Then apply the final `topK(limit)`, ordered by score descending and then row id ascending. Each split only needs `limit` candidates (apache/paimon#9955). - Pin the live-row computation ({{S4}}) and all reads to the plan snapshot (apache/paimon#9031). 5. **Tests.** Cover: - multiple shards - overlapping index ranges - a text column that is an extra field of another index - partition filters - snapshot pinning - deletion-vector tables - a split serialization round trip ### Anything else? - Depends on {{S1}}, {{S2}}, {{S3}} and {{S4}}. - Java's `IndexFullTextSearchSplit` uses Java serialization, so there is no cross-language wire format to match. C++ should define its own versioned serialization. - Java rejects search on tables under query authorization (apache/paimon#8570). C++ should do the same if it adds query auth. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
