SteNicholas opened a new issue, #400: URL: https://github.com/apache/paimon-cpp/issues/400
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation Sub-issue of #399 (step 1: index file format and engine). Paimon C++ builds full-text global indexes with `crates/tantivy_ffi` (tantivy 0.22, jieba-rs 0.7) and its own archive layout. Java Paimon now writes and reads full-text indexes through the standalone engine [apache/paimon-full-text](https://github.com/apache/paimon-full-text) 0.1.0 (`org.apache.paimon:paimon-full-text-index`, apache/paimon#8463, apache/paimon#8467, apache/paimon#8798). As a result, neither side can read the other's index files: | | Java (`paimon-full-text`) | Paimon C++ (`src/paimon/global_index/tantivy/`) | |---|---|---| | Index type / factory | `full-text` (`NativeFullTextGlobalIndexerFactory`) | `tantivy-fulltext` / `tantivy-fulltext-global` | | Option prefix | `full-text.`, stripped once; the remaining keys go to the engine | `tantivy-fulltext.` | | File name prefix | `full-text` | `tantivy-fulltext` | | Archive layout | `PFTIDX01` magic + `u32` version + `u32` header length + JSON header (analyzer config, document count, tantivy version, file table) + concatenated tantivy files | `[i32 count \| (i32 name_len, name, i64 data_len, data)*]` | | Engine | tantivy 0.26; the reader rejects archives whose recorded tantivy version differs. jieba via tantivy-jieba / jieba-rs 0.10.1 | tantivy 0.22, jieba-rs 0.7 + custom `paimon_jieba` that requires `PAIMON_JIEBA_DICT_DIR` | | Analyzer config at read time | Embedded in the archive header; the reader only takes a positional-read callback | Read from `GlobalIndexIOMeta::metadata` and passed to the FFI by the reader | | `GlobalIndexIOMeta` metadata | Flat JSON of the prefix-stripped options (`NativeFullTextIndexOptions#serialize`) | JSON of all prefix-stripped `tantivy-fulltext.*` options | | Writer with no rows | Writes no file (`finish()` returns an empty list) | Always writes an archive (`TantivyGlobalIndexWriter::Finish`) | Consequences today: - `GlobalIndexerFactory::Get("full-text")` finds no `full-text-global` factory and returns `nullptr` (`src/paimon/common/global_index/global_indexer_factory.cpp:53-58`). `GlobalIndexScanImpl::CreateReaders` therefore skips Java-written full-text indexes silently (`src/paimon/core/global_index/global_index_scan_impl.cpp:181-185`), and `GlobalIndexWriteTask` rejects `full-text` as an unknown index type. - `ArchiveLayout::Parse` reads the `PFTI` magic as the file count. The `paimon-ftindex` 0.1.0 reader rejects both C++ fixtures (`test/test_data/cpp_tantivy_fixtures/english_default.archive` and `test/test_data/java_tantivy_fixtures/english_simple.archive`) with `invalid storage format: bad magic`. - The analyzer options differ. C++ supports `tantivy.write.tokenizer`, `jieba.tokenize-mode` and `write.omit-term-freq-and-position`. The `default` tokenizer also behaves differently: Java applies lower-casing, English stemming, stop-word removal and ASCII folding, so a `run` query matches `running`. C++ uses tantivy's `SimpleTokenizer` with lower-casing only. ### Solution 1. **Engine and build** - Replace `crates/tantivy_ffi` with the C API of apache/paimon-full-text pinned to `v0.1.0` (`include/paimon_ftindex.h`). The `paimon-ftindex-ffi` crate is not published on crates.io; only `paimon-ftindex-core` 0.1.0 is. Either import the FFI crate by git tag through Corrosion, or wrap `paimon-ftindex-core` in a thin in-repo static library that exports the same C API. - Keep the tantivy version recorded by Java archives (0.26.x). The reader rejects a mismatched `metadata.tantivy_version`. - Rename the experimental build switch and module accordingly, for example `PAIMON_ENABLE_TANTIVY` → `PAIMON_ENABLE_FULL_TEXT`. The engine tokenizes jieba itself, so the `full-text` path should no longer depend on `PAIMON_JIEBA_DICT_DIR` or on the jieba dictionary from `ThirdpartyToolchain.cmake`, which `lucene-fts` still uses. 2. **Indexer**: register a `full-text` indexer with factory identifier `full-text-global`. - Strip the `full-text.` prefix once and pass the remaining keys to `paimon_ftindex_writer_open(keys, values, len)`. - Supported options and engine defaults: - `tokenizer`: `default`, `simple`, `whitespace`, `raw`, `ngram` or `jieba`; default `default` - `ngram.min-gram`: 3 - `ngram.max-gram`: 3 - `ngram.prefix-only`: false - `jieba.search-mode`: true - `jieba.ordinal-position`: true - `lower-case`: true - `max-token-length`: 40 - `ascii-folding`: true - `stem`: true - `language`: `english` - `remove-stop-words`: true - `stop-words`: empty, `;`-separated - `with-position`: true - Store the prefix-stripped options as flat JSON in `GlobalIndexIOMeta::metadata`, as Java `NativeFullTextIndexOptions#serialize` does. The reader must not depend on this metadata, because the analyzer config is embedded in the archive. 3. **Writer**, following Java `NativeFullTextGlobalIndexWriter`: - Accept `STRING`/`CHAR`/`VARCHAR`. - Add each non-null value with `paimon_ftindex_writer_add_document(writer, relative_row_id, text)`. A null value advances the row count but is not added. - On finish: - If no row was written, return no file. - Otherwise, allocate a `full-text`-prefixed file through `GlobalIndexFileWriter`. - Stream the archive with `paimon_ftindex_writer_write_index` and `PaimonFtindexOutputFile` callbacks. - Return one entry with the row count and the options JSON. 4. **Reader**, following Java `NativeFullTextGlobalIndexReader`: - Require exactly one file per shard. - Open the file lazily with `paimon_ftindex_reader_open`. The `PaimonFtindexInputFile::pread_fn` callback must be safe for concurrent calls; Java serializes `seek` + `readFully` on the input stream. - In `VisitFullTextSearch`, return an empty result for an empty include set. Otherwise: - Call `paimon_ftindex_reader_search`, or `paimon_ftindex_reader_search_with_roaring_filter` with the include row ids serialized as a portable 64-bit Roaring bitmap (the `RoaringTreemap` format). - Size the output buffers to `limit`. - Return a scored result that maps row id to score. - Surface `paimon_ftindex_last_error()` in the returned `Status`. Optionally expose `paimon_ftindex_reader_prewarm` and `paimon_ftindex_reader_read_metrics`. - All other predicate visits are unsupported, as in Java. 5. **Cross-read tests and docs** - Replace the tantivy fixtures with archives generated by the Java `paimon-full-text` module. Cover at least the `default`, `ngram` and `jieba` tokenizers and a null row. Add golden expected row ids and scores for `match`, `match` with `"operator": "And"`, `match_phrase`, `boolean`, and a search with a Roaring filter. - Verify both directions: C++ reads the Java fixtures, and archives written by C++ carry the `PFTIDX01` header and are readable by the Java/Python reader. Tests must not rewrite checked-in fixtures. - Update `docs/source/user_guide/global_index.rst` and `docs/source/building.rst` for the new index type, options and build switch. ### Anything else? - The engine only accepts the JSON DSL. This step can pass `FullTextSearch::query` to the engine unchanged. Aligning the `FullTextSearch` API itself is tracked in {{S2}}. - `tantivy-fulltext` is experimental and off by default. This step replaces it rather than linking two tantivy runtimes, so existing `tantivy-fulltext` indexes must be rebuilt as `full-text`. The `lucene-fts` backend is unaffected. - Java references: - Integration layer: `paimon-full-text/src/main/java/org/apache/paimon/fulltext/index/`. - Engine format: `docs/storage-format.md` in apache/paimon-full-text. Note that `paimon-full-text/README.md` in apache/paimon still describes the old count-prefixed archive layout. - The `paimon-ftindex` 0.1.0 wheel on PyPI wraps the same native engine and can be used to cross-check archives. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
