zhuxiangyi opened a new issue, #962:
URL: https://github.com/apache/paimon-rust/issues/962
### Search before asking
- [x] I searched in the issues and found nothing similar.
### Motivation
paimon-rust can already search `full-text` global indexes
(`FullTextSearchBuilder`,
`full_text_search`, hybrid search), but it cannot build or drop them:
- `CALL sys.create_global_index(..., index_type => 'full-text')` is
rejected. The procedure supports
only `btree`, `bitmap`, `multivalue`, `fm`, and the vindex types.
- `CALL sys.drop_global_index(..., index_type => 'full-text')` is rejected
as an unsupported type.
So a table written by Rust, pypaimon native, or C FFI needs a separate Java
(Flink/Spark) job before
full-text search can use an index. Until then, every search either returns
nothing (`fast` mode)
or builds a temporary in-memory index over the raw rows (`full`/`detail`
mode).
### Solution
Add a native full-text index build that mirrors Java's generic global-index
build
(`GenericIndexTopoBuilder` driving `NativeFullTextGlobalIndexWriter`):
- Split the latest snapshot's row IDs into
`global-index.row-count-per-shard` shards. Skip ranges a
`full-text` index on the column already covers, so repeated builds are
incremental.
- Feed each shard's text column to `paimon-ftindex-core` with shard-relative
row IDs. This is the
native core the Rust reader already uses. Write one `full-text` index file
per non-empty shard
and commit all shards in one snapshot.
- Follow Java's semantics:
- only `full-text.*` options reach the native writer, with the prefix
removed;
- `index_meta` is the flat JSON of those options;
- NULL rows count toward the row count but are not indexed.
- Expose it as `Table::new_full_text_index_build_builder()` (behind the
`fulltext` feature). Route
`create_global_index(index_type => 'full-text')` to it in DataFusion, and
accept `full-text` in
`drop_global_index`.
Table requirements would match the existing global-index builders: row
tracking, data evolution,
global index enabled, no primary keys, no deletion vectors. The column must
be CHAR/VARCHAR.
### Anything else?
Out of scope for the first PR, as possible follow-ups:
- allow tables with deletion vectors (Java's generic build does not reject
them);
- refresh indexes on column updates (`global-index.column-update-action`).
### Willingness to contribute
- [x] I'm willing to submit a PR!
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]