zhuxiangyi opened a new pull request, #963:
URL: https://github.com/apache/paimon-rust/pull/963
### Purpose
Linked issue: close #962
Rust can search `full-text` global indexes (`full_text_search`, hybrid
search), but it cannot
build them: `CALL sys.create_global_index(..., index_type => 'full-text')`
is rejected, and
`drop_global_index` does not accept `full-text` either. Today, tables
written by Rust, pypaimon
native, or C FFI need a Java job to create these indexes.
This PR adds a native full-text index build and the matching drop. It
mirrors Java's generic
global-index build (`GenericIndexTopoBuilder` +
`NativeFullTextGlobalIndexWriter`) and writes the
archive through `paimon-ftindex-core`, the native core that the Rust reader
already uses.
### Brief change log
- Add `FullTextIndexBuildBuilder`
(`Table::new_full_text_index_build_builder`, behind the
`fulltext` feature):
- Splits the latest snapshot's row IDs into
`global-index.row-count-per-shard` shards. It reuses
the existing contiguous shard planner and skips row ranges that a
`full-text` index on the
column already covers, so repeated calls are incremental.
- Feeds each shard's text to `FullTextIndexWriter` with shard-relative row
IDs. Uploads one
`full-text` index file per non-empty shard and commits all shards in one
snapshot (the commit
requires that the snapshot is still the latest). If a shard fails, files
already written are
aborted.
- Matches Java behavior:
- Only `full-text.*` options reach the native writer, with the prefix
removed.
- `index_meta` is the flat JSON of those options.
- NULL rows count toward `row_count` but are not indexed. An all-NULL
shard still writes a
file, so its range is marked as indexed.
- Row IDs must be non-decreasing.
- Shard size is read from the table options merged with the call options.
- Invalid native options (for example an unknown tokenizer) fail before
any data is read.
- Table requirements are the same as the other global-index builders: row
tracking, data
evolution, global index enabled, no primary keys, no deletion vectors.
The column must be
CHAR/VARCHAR. Primary-key tables get a hint to use
`pk-full-text.index.columns`.
- `drop_global_index` accepts `full-text` (case-insensitive) and removes
only the full-text entries
on the requested column.
- DataFusion `create_global_index` routes `index_type => 'full-text'` to the
new builder. Without
the `fulltext` feature it returns a clear "requires the 'fulltext'
feature" error.
- Move `copy_local_file_to_output` from the Lumina writer to
`global_index_build_common` so both
native builders share it. Lumina behavior is unchanged.
### Tests
- `table::full_text_index_build_builder::tests` (13 tests):
- end-to-end build over multiple shards, checking the manifest entries,
row counts including
NULL, the index files, and `FullTextSearchBuilder` results;
- an incremental rebuild that indexes only new rows, and a no-op when
nothing is new;
- an all-NULL shard;
- native options take effect (`full-text.stem=false` changes matching) and
are recorded in
`index_meta`;
- early rejection of invalid options;
- rejection of unsupported tables and columns;
- reading shard rows: out-of-range rows, decreasing row IDs, a NULL
`_ROW_ID`, and
Utf8/LargeUtf8 versus non-string columns.
- `global_index_drop_builder` / `global_index_types`: dropping full-text
removes only the target
field's entries; `full-text` is normalized case-insensitively.
- DataFusion `tests/procedures.rs` (with `fulltext`):
- `test_create_and_drop_full_text_global_index` runs create →
`$table_indexes` →
`full_text_search` → rebuild no-op → drop → search returns nothing;
- a string-column check.
- Existing tests that used `full-text` as their example of an unsupported
type now use types that
are still unsupported.
- `cargo test -p paimon --all-targets --features fulltext,vortex` and `cargo
clippy --all-targets
--workspace --features fulltext,vortex -- -D warnings` pass.
Note: CI runs `paimon-datafusion` tests without the `fulltext` feature, so
the new SQL test is
compiled by clippy but not executed in CI. The existing `full_text_search`
SQL tests have the same
gap. I ran it locally, and I can add a CI step in a follow-up if that's
useful.
### API and Format
- New public API: `Table::new_full_text_index_build_builder()` /
`FullTextIndexBuildBuilder`
(`fulltext` feature).
- `SUPPORTED_GLOBAL_INDEX_TYPES_FOR_DROP` now includes `full-text`.
- No new storage format: the files use the existing `full-text` global index
type and archive
format produced by `paimon-ftindex-core`.
### Documentation
`docs/src/sql.md`:
- documents `index_type => 'full-text'` and its options under
`create_global_index`;
- adds `full-text` to the `drop_global_index` types;
- the Full-Text Search section now explains how to build the index.
Possible follow-ups, not in this PR: allow tables with deletion vectors
(Java does not reject
them), and refresh indexes on column updates
(`global-index.column-update-action`).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]