SteNicholas opened a new issue, #408: URL: https://github.com/apache/paimon-cpp/issues/408
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Motivation Sub-issue of #399 (step 5: primary-key full-text index, options and validation). Paimon C++ recognizes `pk-full-text.index.columns`, but two parts are missing: - **Options are dropped.** `PrimaryKeyIndexDefinitions::Create` builds the `FULL_TEXT` definition with an empty options map (`src/paimon/core/index/pk/primary_key_index_definitions.cpp:208-211`). Tokenizer options such as `full-text.tokenizer=jieba` never reach the index. - **Validation is skipped.** `SchemaValidation::ValidatePrimaryKeyBTreeIndexes` returns early when `pk-btree.index.columns` is empty (`src/paimon/core/schema/schema_validation.cpp:565-570`). A table that sets only `pk-full-text.index.columns` therefore gets none of the checks: column existence, column type, cross-family ownership, deletion vectors, bucket mode. - `docs/source/user_guide/primary_key_global_index.rst` says vector and full-text definitions "are recognized for validation", which is only true when BTree columns are also set. Java behavior (apache/paimon#8651, apache/paimon#8672, apache/paimon#8922): **Option merging** (`CoreOptions#primaryKeyFullTextIndexOptions(column)`): 1. Start from the table options whose key starts with `full-text.`. 2. Parse `fields.<column>.pk-full-text.index.options` as a JSON object of string values. Reject: - anything that is not a JSON object: `<key> must be a JSON object of option key-value pairs.` - an empty key: `<key> contains an empty option key.` - a null value: `<key> value for key <k> must not be null.` 3. Qualify keys that lack the `full-text.` prefix. 4. Reject a key already set at table level with a different value: `<key> defines conflicting values for full-text.<k>.` An equal value is accepted. For example, `{"full-text.tokenizer":"jieba","ngram.min-gram":"2"}` resolves to `full-text.tokenizer=jieba` and `full-text.ngram.min-gram=2`. The `full-text` indexer later strips the prefix (see {{S1}}). **Schema validation** (`SchemaValidation#validatePrimaryKeyFullTextIndex`, run when the key is present): 1. Exactly one column: `pk-full-text.index.columns must contain exactly one column in the first release, but is [...]`. 2. The column is not blank. 3. The table is a primary-key table. 4. `deletion-vectors.enabled = true`, unless the merge engine is `first-row`. 5. `deletion-vectors.merge-on-read = false` when deletion vectors are enabled. 6. Fixed or postpone bucket mode: `bucket > 0` or `bucket = -2`. 7. No `pk-clustering-override`. 8. The column exists. 9. The column type is `CHAR`/`VARCHAR`/`STRING`. 10. The resolved options are valid, as in the merging rules above. Cross-family checks apply to every primary-key index family: - no duplicate column within one family key; - a column can own at most one primary-key index across all families. The maintainer factory also allows only one `FULL_TEXT` definition: `Only one primary-key full-text index is supported.` ### Solution - Resolve the full-text options for each `FULL_TEXT` definition from `fields.<column>.pk-full-text.index.options` and the table-level `full-text.*` options, using the rules above. - Validate primary-key index columns independently of whether BTree columns are set. Add the full-text rules and the cross-family checks. - Fix the statement in `primary_key_global_index.rst`. - Add tests aligned with Java `PrimaryKeyFullTextIndexValidationTest` and `PrimaryKeyIndexDefinitionsTest#testResolvesFullTextIndexOptions`: - more than one column, a duplicate column, a blank column, an unknown column, an `INT` column - deletion vectors off, and `first-row` with deletion vectors off - merge-on-read on - `bucket = -1` rejected, `bucket = -2` accepted - `pk-clustering-override` - the column already used by `pk-btree` - malformed JSON - a conflicting `tokenizer` - an append table ### Anything else? - Java also rejects renaming, dropping, or changing the type of a primary-key index column during schema evolution (`SchemaManagerUtils#assertNotUpdatingPrimaryKeyIndexColumn`). Paimon C++ has no schema-evolution API yet, so this can be added together with one. - This issue unblocks {{S10}} and {{S11}}. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
