SteNicholas opened a new issue, #399: URL: https://github.com/apache/paimon-cpp/issues/399
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation The full-text global index in Paimon C++ (`src/paimon/global_index/tantivy/`, `crates/tantivy_ffi/`) follows the removed Java `paimon-tantivy` module; the Java cross-read fixtures in `test/test_data/java_tantivy_fixtures` were generated on 2026-04-23. Since then Java Paimon has replaced that module with the standalone native engine `paimon-full-text-index` 0.1.0 (apache/paimon#8074, apache/paimon#8308, apache/paimon#8467) and built table-level and primary-key full-text search on top of it. As a result, full-text index files are no longer interoperable between Java and C++ in either direction, and most search features are missing in C++. #### Index file format and engine | | Java (`paimon-full-text-index` 0.1.0) | Paimon C++ | |---|---|---| | Archive layout | `PFTIDX01` magic + version + JSON header (analyzer config and file table) + payloads | Big-endian `[i32 count \| (i32 name_len, name, i64 data_len, data)*]` | | Engine | tantivy 0.26.1 (index format v7), jieba-rs 0.10.1 + tantivy-jieba | tantivy 0.22, jieba-rs 0.7 + custom `paimon_jieba` that requires `PAIMON_JIEBA_DICT_DIR` | | Analyzer config | Embedded in the archive header; the reader only takes a pread callback | Stored in `GlobalIndexIOMeta` metadata and passed to FFI by the reader | | Index type / file prefix | `full-text` / `full-text` | `tantivy-fulltext` (factory `tantivy-fulltext-global`) / `tantivy-fulltext` | | Option prefix | `full-text.` | `tantivy-fulltext.` | | Writer with no rows | Writes no file | Always writes a file | - The `paimon-ftindex` 0.1.0 reader rejects both `test/test_data/cpp_tantivy_fixtures/english_default.archive` and `test/test_data/java_tantivy_fixtures/english_simple.archive` with `invalid storage format: bad magic`. - `ArchiveLayout::Parse` reads the `PFTI` magic as the file count, so C++ cannot read Java archives either. Because no indexer is registered as `full-text`, `GlobalIndexScanImpl` skips Java data-evolution full-text indexes silently. #### Analyzer options - Java supports `full-text.tokenizer` (`default`, `simple`, `whitespace`, `raw`, `ngram`, `jieba`), `ngram.min-gram`, `ngram.max-gram`, `ngram.prefix-only`, `jieba.search-mode`, `jieba.ordinal-position`, `lower-case`, `max-token-length`, `ascii-folding`, `stem`, `language`, `remove-stop-words`, `stop-words` and `with-position`. - C++ supports `tantivy.write.tokenizer` (`default`, `paimon_jieba`, `whitespace`, `raw`, `en_stem`), `jieba.tokenize-mode` and `write.omit-term-freq-and-position`. - The `default` tokenizer behaves differently. Java applies lower-casing, English stemming, stop-word removal and ASCII folding, so a `run` query matches `running`. C++ uses tantivy's `SimpleTokenizer` with lower-casing only. #### Query model (`include/paimon/predicate/full_text_search.h`) - Java takes a JSON DSL query: `match` (operator, boost, fuzziness, prefix_length), `match_phrase` (slop), `boolean` (must/should/must_not), `multi_match` and `boost` (negative_boost). The limit is required and positive, and results are always scored. - C++ takes a `SearchType` enum (`MATCH_ALL`, `MATCH_ANY`, `PHRASE`, `PREFIX`, `WILDCARD`), an optional limit, `with_score` (default `false`) and `min_score`. #### Data-evolution table search C++ only exposes the reader primitive (`GlobalIndexScan::CreateReaders` -> `VisitFullTextSearch` -> `SetGlobalIndexResult`). Compared with Java: - `UnionGlobalIndexReader::VisitFullTextSearch` ORs shard results without a final top-k, so `limit = k` can return up to shards x k rows. Java merges shard results and applies top-k again (apache/paimon#9955). - Deleted rows are not excluded before ranking (apache/paimon#8459, apache/paimon#9031), so they can take top-k slots and the query returns fewer than `limit` rows. - `OffsetGlobalIndexReader::VisitFullTextSearch` shifts every `pre_filter` row id by the shard offset without clipping it to the shard range. - Missing features: - `FullTextSearchBuilder`, `FullTextScan`, `FullTextRead` and serializable search splits. - Row filters via `withFilter`, resolved into include row ids before ranking (apache/paimon#9855). - `full-text-index.search-mode` and `global-index.search-mode`, including `full`/`detail` modes that search unindexed row ranges through a temporary index (apache/paimon#8316, apache/paimon#8844), and `global-index.filter.refine-from-data`. - Range selection across overlapping and multi-field indexes (apache/paimon#8547). - Hybrid search with `rrf`, `weighted_score` and `mrr` rankers (apache/paimon#8271, apache/paimon#8288, apache/paimon#8294, apache/paimon#8348, apache/paimon#8351). - `supportsFullTextSearch()` on indexer factories. #### Primary-key full-text index (`pk-full-text.index.columns`) Java: apache/paimon#8651, apache/paimon#8652, apache/paimon#8672, apache/paimon#9184. - `PrimaryKeyIndexDefinitions` creates the `FULL_TEXT` definition with an empty options map. Java merges table-level `full-text.*` options with `fields.<column>.pk-full-text.index.options` and rejects conflicting values. - Schema validation is missing: exactly one column, primary-key table, deletion vectors enabled, supported bucket mode, `CHAR`/`VARCHAR` type and valid options JSON. - The write side is missing: per-level archive building (`PkFullTextIndexFile`), bucket index state, and `BucketedFullTextIndexMaintainer` with atomic archive replacement on compaction. `BucketedPrimaryKeyIndexMaintainer` only keeps BTree definitions. - The read side is missing: `PrimaryKeyFullTextScan`, `PrimaryKeyFullTextSearchSplit`, `PrimaryKeyFullTextBucketSearch` (deletion vectors, global top-k, parallel search) and `PrimaryKeyFullTextRead`. `PrimaryKeySortedIndexScan` rejects full-text search. - When C++ compacts a table that has Java-written primary-key full-text archives, the archives are neither rebuilt nor retired. ### Solution 1. Replace `crates/tantivy_ffi` with the C API of `paimon-full-text-index` (`paimon_ftindex_writer_open` with key/value options, `paimon_ftindex_writer_add_document`, `paimon_ftindex_writer_write_index` with output callbacks, `paimon_ftindex_reader_open` with a pread callback, `paimon_ftindex_reader_search`, `paimon_ftindex_reader_search_with_roaring_filter`, `paimon_ftindex_reader_prewarm` and `paimon_ftindex_reader_read_metrics`). Register the indexer as `full-text`, use the `full-text.` option and file prefixes, and store the options as flat JSON metadata. Add Java <-> C++ cross-read tests with fixtures generated by the Java `paimon-full-text` module. 2. Align `FullTextSearch` with the JSON DSL query, a required positive limit and scored results. Decide whether `SearchType` stays as a compatibility layer, since the C++-only `lucene-fts` backend uses it. 3. Fix search correctness: apply a final top-k after the union, exclude deleted rows before ranking, and clip `pre_filter` to the shard range. 4. Add table-level full-text search: builder, scan, read and splits, row filters, search modes and hybrid search rankers. 5. Add primary-key full-text index support: options merging and schema validation first, then write-side maintenance and read-side bucket search. Each step can be delivered as a separate PR, and sub-issues can be split from this one. ### Anything else? The format findings were checked with the `paimon-ftindex` 0.1.0 wheel from PyPI, which wraps the same native engine as the Java `paimon-full-text` module. An index written by it has the `PFTIDX01` header described above, and its reader rejects both C++ fixture archives. The other items come from comparing the Java `paimon-full-text` and `paimon-core` sources with the C++ sources. ### Are you willing to submit a PR? - [ ] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
