SteNicholas opened a new issue, #399:
URL: https://github.com/apache/paimon-cpp/issues/399

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   ### Motivation
   
   The full-text global index in Paimon C++ 
(`src/paimon/global_index/tantivy/`, `crates/tantivy_ffi/`) follows the removed 
Java `paimon-tantivy` module; the Java cross-read fixtures in 
`test/test_data/java_tantivy_fixtures` were generated on 2026-04-23. Since then 
Java Paimon has replaced that module with the standalone native engine 
`paimon-full-text-index` 0.1.0 (apache/paimon#8074, apache/paimon#8308, 
apache/paimon#8467) and built table-level and primary-key full-text search on 
top of it. As a result, full-text index files are no longer interoperable 
between Java and C++ in either direction, and most search features are missing 
in C++.
   
   #### Index file format and engine
   
   | | Java (`paimon-full-text-index` 0.1.0) | Paimon C++ |
   |---|---|---|
   | Archive layout | `PFTIDX01` magic + version + JSON header (analyzer config 
and file table) + payloads | Big-endian `[i32 count \| (i32 name_len, name, i64 
data_len, data)*]` |
   | Engine | tantivy 0.26.1 (index format v7), jieba-rs 0.10.1 + tantivy-jieba 
| tantivy 0.22, jieba-rs 0.7 + custom `paimon_jieba` that requires 
`PAIMON_JIEBA_DICT_DIR` |
   | Analyzer config | Embedded in the archive header; the reader only takes a 
pread callback | Stored in `GlobalIndexIOMeta` metadata and passed to FFI by 
the reader |
   | Index type / file prefix | `full-text` / `full-text` | `tantivy-fulltext` 
(factory `tantivy-fulltext-global`) / `tantivy-fulltext` |
   | Option prefix | `full-text.` | `tantivy-fulltext.` |
   | Writer with no rows | Writes no file | Always writes a file |
   
   - The `paimon-ftindex` 0.1.0 reader rejects both 
`test/test_data/cpp_tantivy_fixtures/english_default.archive` and 
`test/test_data/java_tantivy_fixtures/english_simple.archive` with `invalid 
storage format: bad magic`.
   - `ArchiveLayout::Parse` reads the `PFTI` magic as the file count, so C++ 
cannot read Java archives either. Because no indexer is registered as 
`full-text`, `GlobalIndexScanImpl` skips Java data-evolution full-text indexes 
silently.
   
   #### Analyzer options
   
   - Java supports `full-text.tokenizer` (`default`, `simple`, `whitespace`, 
`raw`, `ngram`, `jieba`), `ngram.min-gram`, `ngram.max-gram`, 
`ngram.prefix-only`, `jieba.search-mode`, `jieba.ordinal-position`, 
`lower-case`, `max-token-length`, `ascii-folding`, `stem`, `language`, 
`remove-stop-words`, `stop-words` and `with-position`.
   - C++ supports `tantivy.write.tokenizer` (`default`, `paimon_jieba`, 
`whitespace`, `raw`, `en_stem`), `jieba.tokenize-mode` and 
`write.omit-term-freq-and-position`.
   - The `default` tokenizer behaves differently. Java applies lower-casing, 
English stemming, stop-word removal and ASCII folding, so a `run` query matches 
`running`. C++ uses tantivy's `SimpleTokenizer` with lower-casing only.
   
   #### Query model (`include/paimon/predicate/full_text_search.h`)
   
   - Java takes a JSON DSL query: `match` (operator, boost, fuzziness, 
prefix_length), `match_phrase` (slop), `boolean` (must/should/must_not), 
`multi_match` and `boost` (negative_boost). The limit is required and positive, 
and results are always scored.
   - C++ takes a `SearchType` enum (`MATCH_ALL`, `MATCH_ANY`, `PHRASE`, 
`PREFIX`, `WILDCARD`), an optional limit, `with_score` (default `false`) and 
`min_score`.
   
   #### Data-evolution table search
   
   C++ only exposes the reader primitive (`GlobalIndexScan::CreateReaders` -> 
`VisitFullTextSearch` -> `SetGlobalIndexResult`). Compared with Java:
   
   - `UnionGlobalIndexReader::VisitFullTextSearch` ORs shard results without a 
final top-k, so `limit = k` can return up to shards x k rows. Java merges shard 
results and applies top-k again (apache/paimon#9955).
   - Deleted rows are not excluded before ranking (apache/paimon#8459, 
apache/paimon#9031), so they can take top-k slots and the query returns fewer 
than `limit` rows.
   - `OffsetGlobalIndexReader::VisitFullTextSearch` shifts every `pre_filter` 
row id by the shard offset without clipping it to the shard range.
   - Missing features:
     - `FullTextSearchBuilder`, `FullTextScan`, `FullTextRead` and serializable 
search splits.
     - Row filters via `withFilter`, resolved into include row ids before 
ranking (apache/paimon#9855).
     - `full-text-index.search-mode` and `global-index.search-mode`, including 
`full`/`detail` modes that search unindexed row ranges through a temporary 
index (apache/paimon#8316, apache/paimon#8844), and 
`global-index.filter.refine-from-data`.
     - Range selection across overlapping and multi-field indexes 
(apache/paimon#8547).
     - Hybrid search with `rrf`, `weighted_score` and `mrr` rankers 
(apache/paimon#8271, apache/paimon#8288, apache/paimon#8294, 
apache/paimon#8348, apache/paimon#8351).
     - `supportsFullTextSearch()` on indexer factories.
   
   #### Primary-key full-text index (`pk-full-text.index.columns`)
   
   Java: apache/paimon#8651, apache/paimon#8652, apache/paimon#8672, 
apache/paimon#9184.
   
   - `PrimaryKeyIndexDefinitions` creates the `FULL_TEXT` definition with an 
empty options map. Java merges table-level `full-text.*` options with 
`fields.<column>.pk-full-text.index.options` and rejects conflicting values.
   - Schema validation is missing: exactly one column, primary-key table, 
deletion vectors enabled, supported bucket mode, `CHAR`/`VARCHAR` type and 
valid options JSON.
   - The write side is missing: per-level archive building 
(`PkFullTextIndexFile`), bucket index state, and 
`BucketedFullTextIndexMaintainer` with atomic archive replacement on 
compaction. `BucketedPrimaryKeyIndexMaintainer` only keeps BTree definitions.
   - The read side is missing: `PrimaryKeyFullTextScan`, 
`PrimaryKeyFullTextSearchSplit`, `PrimaryKeyFullTextBucketSearch` (deletion 
vectors, global top-k, parallel search) and `PrimaryKeyFullTextRead`. 
`PrimaryKeySortedIndexScan` rejects full-text search.
   - When C++ compacts a table that has Java-written primary-key full-text 
archives, the archives are neither rebuilt nor retired.
   
   ### Solution
   
   1. Replace `crates/tantivy_ffi` with the C API of `paimon-full-text-index` 
(`paimon_ftindex_writer_open` with key/value options, 
`paimon_ftindex_writer_add_document`, `paimon_ftindex_writer_write_index` with 
output callbacks, `paimon_ftindex_reader_open` with a pread callback, 
`paimon_ftindex_reader_search`, 
`paimon_ftindex_reader_search_with_roaring_filter`, 
`paimon_ftindex_reader_prewarm` and `paimon_ftindex_reader_read_metrics`). 
Register the indexer as `full-text`, use the `full-text.` option and file 
prefixes, and store the options as flat JSON metadata. Add Java <-> C++ 
cross-read tests with fixtures generated by the Java `paimon-full-text` module.
   2. Align `FullTextSearch` with the JSON DSL query, a required positive limit 
and scored results. Decide whether `SearchType` stays as a compatibility layer, 
since the C++-only `lucene-fts` backend uses it.
   3. Fix search correctness: apply a final top-k after the union, exclude 
deleted rows before ranking, and clip `pre_filter` to the shard range.
   4. Add table-level full-text search: builder, scan, read and splits, row 
filters, search modes and hybrid search rankers.
   5. Add primary-key full-text index support: options merging and schema 
validation first, then write-side maintenance and read-side bucket search.
   
   Each step can be delivered as a separate PR, and sub-issues can be split 
from this one.
   
   ### Anything else?
   
   The format findings were checked with the `paimon-ftindex` 0.1.0 wheel from 
PyPI, which wraps the same native engine as the Java `paimon-full-text` module. 
An index written by it has the `PFTIDX01` header described above, and its 
reader rejects both C++ fixture archives. The other items come from comparing 
the Java `paimon-full-text` and `paimon-core` sources with the C++ sources.
   
   ### Are you willing to submit a PR?
   
   - [ ] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to