zhuxiangyi opened a new pull request, #963:
URL: https://github.com/apache/paimon-rust/pull/963

   ### Purpose
   
   Linked issue: close #962
   
   Rust can search `full-text` global indexes (`full_text_search`, hybrid 
search), but it cannot
   build them: `CALL sys.create_global_index(..., index_type => 'full-text')` 
is rejected, and
   `drop_global_index` does not accept `full-text` either. Today, tables 
written by Rust, pypaimon
   native, or C FFI need a Java job to create these indexes.
   
   This PR adds a native full-text index build and the matching drop. It 
mirrors Java's generic
   global-index build (`GenericIndexTopoBuilder` + 
`NativeFullTextGlobalIndexWriter`) and writes the
   archive through `paimon-ftindex-core`, the native core that the Rust reader 
already uses.
   
   ### Brief change log
   
   - Add `FullTextIndexBuildBuilder` 
(`Table::new_full_text_index_build_builder`, behind the
     `fulltext` feature):
     - Splits the latest snapshot's row IDs into 
`global-index.row-count-per-shard` shards. It reuses
       the existing contiguous shard planner and skips row ranges that a 
`full-text` index on the
       column already covers, so repeated calls are incremental.
     - Feeds each shard's text to `FullTextIndexWriter` with shard-relative row 
IDs. Uploads one
       `full-text` index file per non-empty shard and commits all shards in one 
snapshot (the commit
       requires that the snapshot is still the latest). If a shard fails, files 
already written are
       aborted.
     - Matches Java behavior:
       - Only `full-text.*` options reach the native writer, with the prefix 
removed.
       - `index_meta` is the flat JSON of those options.
       - NULL rows count toward `row_count` but are not indexed. An all-NULL 
shard still writes a
         file, so its range is marked as indexed.
       - Row IDs must be non-decreasing.
       - Shard size is read from the table options merged with the call options.
     - Invalid native options (for example an unknown tokenizer) fail before 
any data is read.
     - Table requirements are the same as the other global-index builders: row 
tracking, data
       evolution, global index enabled, no primary keys, no deletion vectors. 
The column must be
       CHAR/VARCHAR. Primary-key tables get a hint to use 
`pk-full-text.index.columns`.
   - `drop_global_index` accepts `full-text` (case-insensitive) and removes 
only the full-text entries
     on the requested column.
   - DataFusion `create_global_index` routes `index_type => 'full-text'` to the 
new builder. Without
     the `fulltext` feature it returns a clear "requires the 'fulltext' 
feature" error.
   - Move `copy_local_file_to_output` from the Lumina writer to 
`global_index_build_common` so both
     native builders share it. Lumina behavior is unchanged.
   
   ### Tests
   
   - `table::full_text_index_build_builder::tests` (13 tests):
     - end-to-end build over multiple shards, checking the manifest entries, 
row counts including
       NULL, the index files, and `FullTextSearchBuilder` results;
     - an incremental rebuild that indexes only new rows, and a no-op when 
nothing is new;
     - an all-NULL shard;
     - native options take effect (`full-text.stem=false` changes matching) and 
are recorded in
       `index_meta`;
     - early rejection of invalid options;
     - rejection of unsupported tables and columns;
     - reading shard rows: out-of-range rows, decreasing row IDs, a NULL 
`_ROW_ID`, and
       Utf8/LargeUtf8 versus non-string columns.
   - `global_index_drop_builder` / `global_index_types`: dropping full-text 
removes only the target
     field's entries; `full-text` is normalized case-insensitively.
   - DataFusion `tests/procedures.rs` (with `fulltext`):
     - `test_create_and_drop_full_text_global_index` runs create → 
`$table_indexes` →
       `full_text_search` → rebuild no-op → drop → search returns nothing;
     - a string-column check.
     - Existing tests that used `full-text` as their example of an unsupported 
type now use types that
       are still unsupported.
   - `cargo test -p paimon --all-targets --features fulltext,vortex` and `cargo 
clippy --all-targets
     --workspace --features fulltext,vortex -- -D warnings` pass.
   
   Note: CI runs `paimon-datafusion` tests without the `fulltext` feature, so 
the new SQL test is
   compiled by clippy but not executed in CI. The existing `full_text_search` 
SQL tests have the same
   gap. I ran it locally, and I can add a CI step in a follow-up if that's 
useful.
   
   ### API and Format
   
   - New public API: `Table::new_full_text_index_build_builder()` / 
`FullTextIndexBuildBuilder`
     (`fulltext` feature).
   - `SUPPORTED_GLOBAL_INDEX_TYPES_FOR_DROP` now includes `full-text`.
   - No new storage format: the files use the existing `full-text` global index 
type and archive
     format produced by `paimon-ftindex-core`.
   
   ### Documentation
   
   `docs/src/sql.md`:
   - documents `index_type => 'full-text'` and its options under 
`create_global_index`;
   - adds `full-text` to the `drop_global_index` types;
   - the Full-Text Search section now explains how to build the index.
   
   Possible follow-ups, not in this PR: allow tables with deletion vectors 
(Java does not reject
   them), and refresh indexes on column updates 
(`global-index.column-update-action`).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to