zhuxiangyi opened a new issue, #962:
URL: https://github.com/apache/paimon-rust/issues/962

   ### Search before asking
   
   - [x] I searched in the issues and found nothing similar.
   
   ### Motivation
   
   paimon-rust can already search `full-text` global indexes 
(`FullTextSearchBuilder`,
   `full_text_search`, hybrid search), but it cannot build or drop them:
   
   - `CALL sys.create_global_index(..., index_type => 'full-text')` is 
rejected. The procedure supports
     only `btree`, `bitmap`, `multivalue`, `fm`, and the vindex types.
   - `CALL sys.drop_global_index(..., index_type => 'full-text')` is rejected 
as an unsupported type.
   
   So a table written by Rust, pypaimon native, or C FFI needs a separate Java 
(Flink/Spark) job before
   full-text search can use an index. Until then, every search either returns 
nothing (`fast` mode)
   or builds a temporary in-memory index over the raw rows (`full`/`detail` 
mode).
   
   ### Solution
   
   Add a native full-text index build that mirrors Java's generic global-index 
build
   (`GenericIndexTopoBuilder` driving `NativeFullTextGlobalIndexWriter`):
   
   - Split the latest snapshot's row IDs into 
`global-index.row-count-per-shard` shards. Skip ranges a
     `full-text` index on the column already covers, so repeated builds are 
incremental.
   - Feed each shard's text column to `paimon-ftindex-core` with shard-relative 
row IDs. This is the
     native core the Rust reader already uses. Write one `full-text` index file 
per non-empty shard
     and commit all shards in one snapshot.
   - Follow Java's semantics:
     - only `full-text.*` options reach the native writer, with the prefix 
removed;
     - `index_meta` is the flat JSON of those options;
     - NULL rows count toward the row count but are not indexed.
   - Expose it as `Table::new_full_text_index_build_builder()` (behind the 
`fulltext` feature). Route
     `create_global_index(index_type => 'full-text')` to it in DataFusion, and 
accept `full-text` in
     `drop_global_index`.
   
   Table requirements would match the existing global-index builders: row 
tracking, data evolution,
   global index enabled, no primary keys, no deletion vectors. The column must 
be CHAR/VARCHAR.
   
   ### Anything else?
   
   Out of scope for the first PR, as possible follow-ups:
   
   - allow tables with deletion vectors (Java's generic build does not reject 
them);
   - refresh indexes on column updates (`global-index.column-update-action`).
   
   ### Willingness to contribute
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to