SteNicholas opened a new issue, #408:
URL: https://github.com/apache/paimon-cpp/issues/408

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   ### Motivation
   
   Sub-issue of #399 (step 5: primary-key full-text index, options and 
validation).
   
   Paimon C++ recognizes `pk-full-text.index.columns`, but two parts are 
missing:
   
   - **Options are dropped.** `PrimaryKeyIndexDefinitions::Create` builds the 
`FULL_TEXT` definition with an empty options map 
(`src/paimon/core/index/pk/primary_key_index_definitions.cpp:208-211`). 
Tokenizer options such as `full-text.tokenizer=jieba` never reach the index.
   - **Validation is skipped.** 
`SchemaValidation::ValidatePrimaryKeyBTreeIndexes` returns early when 
`pk-btree.index.columns` is empty 
(`src/paimon/core/schema/schema_validation.cpp:565-570`). A table that sets 
only `pk-full-text.index.columns` therefore gets none of the checks: column 
existence, column type, cross-family ownership, deletion vectors, bucket mode.
     - `docs/source/user_guide/primary_key_global_index.rst` says vector and 
full-text definitions "are recognized for validation", which is only true when 
BTree columns are also set.
   
   Java behavior (apache/paimon#8651, apache/paimon#8672, apache/paimon#8922):
   
   **Option merging** (`CoreOptions#primaryKeyFullTextIndexOptions(column)`):
   
   1. Start from the table options whose key starts with `full-text.`.
   2. Parse `fields.<column>.pk-full-text.index.options` as a JSON object of 
string values. Reject:
      - anything that is not a JSON object: `<key> must be a JSON object of 
option key-value pairs.`
      - an empty key: `<key> contains an empty option key.`
      - a null value: `<key> value for key <k> must not be null.`
   3. Qualify keys that lack the `full-text.` prefix.
   4. Reject a key already set at table level with a different value: `<key> 
defines conflicting values for full-text.<k>.` An equal value is accepted.
   
   For example, `{"full-text.tokenizer":"jieba","ngram.min-gram":"2"}` resolves 
to `full-text.tokenizer=jieba` and `full-text.ngram.min-gram=2`. The 
`full-text` indexer later strips the prefix (see {{S1}}).
   
   **Schema validation** (`SchemaValidation#validatePrimaryKeyFullTextIndex`, 
run when the key is present):
   
   1. Exactly one column: `pk-full-text.index.columns must contain exactly one 
column in the first release, but is [...]`.
   2. The column is not blank.
   3. The table is a primary-key table.
   4. `deletion-vectors.enabled = true`, unless the merge engine is `first-row`.
   5. `deletion-vectors.merge-on-read = false` when deletion vectors are 
enabled.
   6. Fixed or postpone bucket mode: `bucket > 0` or `bucket = -2`.
   7. No `pk-clustering-override`.
   8. The column exists.
   9. The column type is `CHAR`/`VARCHAR`/`STRING`.
   10. The resolved options are valid, as in the merging rules above.
   
   Cross-family checks apply to every primary-key index family:
   - no duplicate column within one family key;
   - a column can own at most one primary-key index across all families.
   
   The maintainer factory also allows only one `FULL_TEXT` definition: `Only 
one primary-key full-text index is supported.`
   
   ### Solution
   
   - Resolve the full-text options for each `FULL_TEXT` definition from 
`fields.<column>.pk-full-text.index.options` and the table-level `full-text.*` 
options, using the rules above.
   - Validate primary-key index columns independently of whether BTree columns 
are set. Add the full-text rules and the cross-family checks.
   - Fix the statement in `primary_key_global_index.rst`.
   - Add tests aligned with Java `PrimaryKeyFullTextIndexValidationTest` and 
`PrimaryKeyIndexDefinitionsTest#testResolvesFullTextIndexOptions`:
     - more than one column, a duplicate column, a blank column, an unknown 
column, an `INT` column
     - deletion vectors off, and `first-row` with deletion vectors off
     - merge-on-read on
     - `bucket = -1` rejected, `bucket = -2` accepted
     - `pk-clustering-override`
     - the column already used by `pk-btree`
     - malformed JSON
     - a conflicting `tokenizer`
     - an append table
   
   ### Anything else?
   
   - Java also rejects renaming, dropping, or changing the type of a 
primary-key index column during schema evolution 
(`SchemaManagerUtils#assertNotUpdatingPrimaryKeyIndexColumn`). Paimon C++ has 
no schema-evolution API yet, so this can be added together with one.
   - This issue unblocks {{S10}} and {{S11}}.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to