SteNicholas opened a new issue, #400:
URL: https://github.com/apache/paimon-cpp/issues/400

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   ### Motivation
   
   Sub-issue of #399 (step 1: index file format and engine).
   
   Paimon C++ builds full-text global indexes with `crates/tantivy_ffi` 
(tantivy 0.22, jieba-rs 0.7) and its own archive layout. Java Paimon now writes 
and reads full-text indexes through the standalone engine 
[apache/paimon-full-text](https://github.com/apache/paimon-full-text) 0.1.0 
(`org.apache.paimon:paimon-full-text-index`, apache/paimon#8463, 
apache/paimon#8467, apache/paimon#8798). As a result, neither side can read the 
other's index files:
   
   | | Java (`paimon-full-text`) | Paimon C++ 
(`src/paimon/global_index/tantivy/`) |
   |---|---|---|
   | Index type / factory | `full-text` (`NativeFullTextGlobalIndexerFactory`) 
| `tantivy-fulltext` / `tantivy-fulltext-global` |
   | Option prefix | `full-text.`, stripped once; the remaining keys go to the 
engine | `tantivy-fulltext.` |
   | File name prefix | `full-text` | `tantivy-fulltext` |
   | Archive layout | `PFTIDX01` magic + `u32` version + `u32` header length + 
JSON header (analyzer config, document count, tantivy version, file table) + 
concatenated tantivy files | `[i32 count \| (i32 name_len, name, i64 data_len, 
data)*]` |
   | Engine | tantivy 0.26; the reader rejects archives whose recorded tantivy 
version differs. jieba via tantivy-jieba / jieba-rs 0.10.1 | tantivy 0.22, 
jieba-rs 0.7 + custom `paimon_jieba` that requires `PAIMON_JIEBA_DICT_DIR` |
   | Analyzer config at read time | Embedded in the archive header; the reader 
only takes a positional-read callback | Read from `GlobalIndexIOMeta::metadata` 
and passed to the FFI by the reader |
   | `GlobalIndexIOMeta` metadata | Flat JSON of the prefix-stripped options 
(`NativeFullTextIndexOptions#serialize`) | JSON of all prefix-stripped 
`tantivy-fulltext.*` options |
   | Writer with no rows | Writes no file (`finish()` returns an empty list) | 
Always writes an archive (`TantivyGlobalIndexWriter::Finish`) |
   
   Consequences today:
   
   - `GlobalIndexerFactory::Get("full-text")` finds no `full-text-global` 
factory and returns `nullptr` 
(`src/paimon/common/global_index/global_indexer_factory.cpp:53-58`). 
`GlobalIndexScanImpl::CreateReaders` therefore skips Java-written full-text 
indexes silently 
(`src/paimon/core/global_index/global_index_scan_impl.cpp:181-185`), and 
`GlobalIndexWriteTask` rejects `full-text` as an unknown index type.
   - `ArchiveLayout::Parse` reads the `PFTI` magic as the file count. The 
`paimon-ftindex` 0.1.0 reader rejects both C++ fixtures 
(`test/test_data/cpp_tantivy_fixtures/english_default.archive` and 
`test/test_data/java_tantivy_fixtures/english_simple.archive`) with `invalid 
storage format: bad magic`.
   - The analyzer options differ. C++ supports `tantivy.write.tokenizer`, 
`jieba.tokenize-mode` and `write.omit-term-freq-and-position`. The `default` 
tokenizer also behaves differently: Java applies lower-casing, English 
stemming, stop-word removal and ASCII folding, so a `run` query matches 
`running`. C++ uses tantivy's `SimpleTokenizer` with lower-casing only.
   
   ### Solution
   
   1. **Engine and build**
      - Replace `crates/tantivy_ffi` with the C API of apache/paimon-full-text 
pinned to `v0.1.0` (`include/paimon_ftindex.h`). The `paimon-ftindex-ffi` crate 
is not published on crates.io; only `paimon-ftindex-core` 0.1.0 is. Either 
import the FFI crate by git tag through Corrosion, or wrap 
`paimon-ftindex-core` in a thin in-repo static library that exports the same C 
API.
      - Keep the tantivy version recorded by Java archives (0.26.x). The reader 
rejects a mismatched `metadata.tantivy_version`.
      - Rename the experimental build switch and module accordingly, for 
example `PAIMON_ENABLE_TANTIVY` → `PAIMON_ENABLE_FULL_TEXT`. The engine 
tokenizes jieba itself, so the `full-text` path should no longer depend on 
`PAIMON_JIEBA_DICT_DIR` or on the jieba dictionary from 
`ThirdpartyToolchain.cmake`, which `lucene-fts` still uses.
   2. **Indexer**: register a `full-text` indexer with factory identifier 
`full-text-global`.
      - Strip the `full-text.` prefix once and pass the remaining keys to 
`paimon_ftindex_writer_open(keys, values, len)`.
      - Supported options and engine defaults:
        - `tokenizer`: `default`, `simple`, `whitespace`, `raw`, `ngram` or 
`jieba`; default `default`
        - `ngram.min-gram`: 3
        - `ngram.max-gram`: 3
        - `ngram.prefix-only`: false
        - `jieba.search-mode`: true
        - `jieba.ordinal-position`: true
        - `lower-case`: true
        - `max-token-length`: 40
        - `ascii-folding`: true
        - `stem`: true
        - `language`: `english`
        - `remove-stop-words`: true
        - `stop-words`: empty, `;`-separated
        - `with-position`: true
      - Store the prefix-stripped options as flat JSON in 
`GlobalIndexIOMeta::metadata`, as Java `NativeFullTextIndexOptions#serialize` 
does. The reader must not depend on this metadata, because the analyzer config 
is embedded in the archive.
   3. **Writer**, following Java `NativeFullTextGlobalIndexWriter`:
      - Accept `STRING`/`CHAR`/`VARCHAR`.
      - Add each non-null value with 
`paimon_ftindex_writer_add_document(writer, relative_row_id, text)`. A null 
value advances the row count but is not added.
      - On finish:
        - If no row was written, return no file.
        - Otherwise, allocate a `full-text`-prefixed file through 
`GlobalIndexFileWriter`.
        - Stream the archive with `paimon_ftindex_writer_write_index` and 
`PaimonFtindexOutputFile` callbacks.
        - Return one entry with the row count and the options JSON.
   4. **Reader**, following Java `NativeFullTextGlobalIndexReader`:
      - Require exactly one file per shard.
      - Open the file lazily with `paimon_ftindex_reader_open`. The 
`PaimonFtindexInputFile::pread_fn` callback must be safe for concurrent calls; 
Java serializes `seek` + `readFully` on the input stream.
      - In `VisitFullTextSearch`, return an empty result for an empty include 
set. Otherwise:
        - Call `paimon_ftindex_reader_search`, or 
`paimon_ftindex_reader_search_with_roaring_filter` with the include row ids 
serialized as a portable 64-bit Roaring bitmap (the `RoaringTreemap` format).
        - Size the output buffers to `limit`.
        - Return a scored result that maps row id to score.
      - Surface `paimon_ftindex_last_error()` in the returned `Status`. 
Optionally expose `paimon_ftindex_reader_prewarm` and 
`paimon_ftindex_reader_read_metrics`.
      - All other predicate visits are unsupported, as in Java.
   5. **Cross-read tests and docs**
      - Replace the tantivy fixtures with archives generated by the Java 
`paimon-full-text` module. Cover at least the `default`, `ngram` and `jieba` 
tokenizers and a null row. Add golden expected row ids and scores for `match`, 
`match` with `"operator": "And"`, `match_phrase`, `boolean`, and a search with 
a Roaring filter.
      - Verify both directions: C++ reads the Java fixtures, and archives 
written by C++ carry the `PFTIDX01` header and are readable by the Java/Python 
reader. Tests must not rewrite checked-in fixtures.
      - Update `docs/source/user_guide/global_index.rst` and 
`docs/source/building.rst` for the new index type, options and build switch.
   
   ### Anything else?
   
   - The engine only accepts the JSON DSL. This step can pass 
`FullTextSearch::query` to the engine unchanged. Aligning the `FullTextSearch` 
API itself is tracked in {{S2}}.
   - `tantivy-fulltext` is experimental and off by default. This step replaces 
it rather than linking two tantivy runtimes, so existing `tantivy-fulltext` 
indexes must be rebuilt as `full-text`. The `lucene-fts` backend is unaffected.
   - Java references:
     - Integration layer: 
`paimon-full-text/src/main/java/org/apache/paimon/fulltext/index/`.
     - Engine format: `docs/storage-format.md` in apache/paimon-full-text. Note 
that `paimon-full-text/README.md` in apache/paimon still describes the old 
count-prefixed archive layout.
   - The `paimon-ftindex` 0.1.0 wheel on PyPI wraps the same native engine and 
can be used to cross-check archives.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to