airborne12 commented on code in PR #67918:
URL: https://github.com/apache/doris/pull/67918#discussion_r4089174214


##########
fe/fe-core/src/main/java/org/apache/doris/analysis/invertedindex/AnalyzerIdentityBuilder.java:
##########
@@ -45,14 +118,96 @@ public static String buildAnalyzerIdentity(
         }
 
         if (!Strings.isNullOrEmpty(preferredAnalyzer)) {
+            String builtinIkIdentity = 
resolveBuiltinIkAnalyzerIdentity(properties, preferredAnalyzer);
+            if (builtinIkIdentity != null) {
+                return appendOuterCharFilterIdentity(
+                        builtinIkIdentity, properties, 
builtinIkFoldContext(builtinIkIdentity));
+            }
+            // BE dispatches a canonical lowercase built-in before any custom 
policy of that name.
+            String builtinIdentity = 
builtinAnalyzerIdentity(preferredAnalyzer.trim(), properties);
+            if (builtinIdentity != null) {
+                return builtinIdentity;
+            }
             // For custom analyzer/normalizer, resolve to underlying config to 
build identity
-            return resolveAnalyzerIdentity(preferredAnalyzer, 
defaultAnalyzerKey, log);
+            return appendOuterCharFilterIdentity(
+                    resolveAnalyzerIdentity(preferredAnalyzer, 
defaultAnalyzerKey, log), properties,
+                    customAnalyzerFoldContext(preferredAnalyzer));
         }
 
         if (Strings.isNullOrEmpty(parser) || 
parserNone.equalsIgnoreCase(parser)) {
             return defaultAnalyzerKey;
         }
-        return parser;
+        String legacyIkIdentity = resolveLegacyIkIdentity(properties, parser);
+        if (legacyIkIdentity != null) {
+            return appendOuterCharFilterIdentity(
+                    legacyIkIdentity, properties, 
builtinIkFoldContext(legacyIkIdentity));
+        }
+        // A parser name reaches BE's built-in dispatch after case folding and 
without a policy lookup.
+        String canonicalParser = parser.trim().toLowerCase(Locale.ROOT);
+        String builtinParserIdentity = 
builtinAnalyzerIdentity(canonicalParser, properties);
+        if (builtinParserIdentity != null) {
+            return builtinParserIdentity;
+        }
+        return appendOuterCharFilterIdentity(canonicalParser, properties, 
null);
+    }
+
+    /**
+     * Identity of a built-in analyzer BE canonicalizes, or null for any other 
name. The basic and
+     * icu built-ins are their tokenizer plus a LowerCaseFilter that 
lower_case drops, so they share
+     * the identity of that custom pipeline, and unicode is another spelling 
of standard.
+     */
+    private static String builtinAnalyzerIdentity(String name, Map<String, 
String> properties) {
+        if 
(InvertedIndexProperties.INVERTED_INDEX_PARSER_UNICODE.equals(name)) {

Review Comment:
   Confirmed, and this one was ours - the previous commit introduced it. Fixed 
in e3a55f591a1.
   
   `standard95::StandardAnalyzer` passes `_lowercase` and `_stopwords` straight 
into `StandardTokenizer` (`StandardAnalyzer.h:19,25`), and the tokenizer acts 
on both: it lower-cases a whole `WORD_TYPE` term and skips a term found in the 
stop word set (`StandardTokenizer.h:50-57`). Collapsing both spellings to a 
bare `standard` therefore did worse than lose a distinction - it made the 
fences reject a pair of indexes that really do emit different terms.
   
   The identity now keeps both settings, appending `lower_case=false` and 
`stopwords=none` only when they differ from what BE applies by default, so the 
two spellings stay equal under equal settings and differ under mixed ones. When 
lower-casing is on, the identity also hands `FoldContext.unfiltered()` to the 
outer char filter, so an outer `A -> a` is absorbed the same way it is on the 
custom side.
   
   The same class of bug was still present on three other built-ins that this 
thread did not name, and we fixed them together: `english` (`SimpleAnalyzer` 
passes `_lowercase` to `SimpleTokenizer`, `Analyzers.h:191,195`), `chinese` 
(`LanguageBasedAnalyzer.cpp:91,124` passes it to `ChineseTokenizer`) and 
`kuromoji` (`KuromojiAnalyzer.h:55,63`). They only need the `lower_case` 
setting, because `_stopwords` is read by `StandardTokenizer` alone.
   
   Tests: `testBuiltinStandardIdentityKeepsEffectiveSettings` for the 
mixed-setting and filtered/unfiltered coverage, and 
`testBuiltinAnalyzersReadingOnlyLowerCaseKeepThatSetting` for the other three.



##########
fe/fe-core/src/main/java/org/apache/doris/analysis/invertedindex/AnalyzerIdentityBuilder.java:
##########
@@ -150,49 +343,592 @@ private static String 
buildIdentityFromPolicyProperties(IndexPolicyTypeEnum type
      * Resolve a component (tokenizer) to its identity.
      */
     private static String resolveComponentIdentity(String name, 
IndexPolicyTypeEnum expectedType) {
+        return resolveComponentIdentity(name, expectedType, null);
+    }
+
+    /** {@code fold} is the case-folding context of a char filter, or null 
without a downstream fold. */
+    private static String resolveComponentIdentity(
+            String name, IndexPolicyTypeEnum expectedType, FoldContext fold) {
         if (Strings.isNullOrEmpty(name)) {
             return "";
         }
 
-        // Check if it's a built-in component
-        if (expectedType == IndexPolicyTypeEnum.TOKENIZER
-                && IndexPolicy.BUILTIN_TOKENIZERS.contains(name)) {
-            return name;
+        // Existing named policies take precedence over built-ins for upgrade 
compatibility.
+        try {
+            Env env = Env.getCurrentEnv();
+            if (env != null && env.getIndexPolicyMgr() != null) {
+                IndexPolicy policy = 
env.getIndexPolicyMgr().getPolicyByName(name);
+                if (policy != null && policy.getType() == expectedType) {
+                    if (policy.isInvalid()) {
+                        return "invalid-policy:" + policy.getId() + ":" + 
policy.getName();
+                    }
+                    Map<String, String> props = policy.getProperties();
+                    if (props != null && !props.isEmpty()) {
+                        TreeMap<String, String> sortedProps = new 
TreeMap<>(props);
+                        String type = sortedProps.get(IndexPolicy.PROP_TYPE);
+                        String normalizedType = 
normalizeBuiltinComponentName(type, expectedType);
+                        if (normalizedType != null) {
+                            if ("empty".equals(normalizedType)) {
+                                return "";
+                            }
+                            sortedProps.put(IndexPolicy.PROP_TYPE, 
normalizedType);
+                            canonicalizeEffectiveComponentProperties(
+                                    sortedProps, normalizedType, expectedType);
+                            if (sortedProps.size() == 1) {
+                                return normalizedType;
+                            }
+                        }
+                        if (expectedType == IndexPolicyTypeEnum.TOKENIZER
+                                && 
"ngram".equals(sortedProps.get(IndexPolicy.PROP_TYPE))) {
+                            // This setting only limits policy creation; it 
does not change emitted tokens.
+                            sortedProps.remove(PROP_MAX_NGRAM_DIFF);
+                        }
+                        if (expectedType == IndexPolicyTypeEnum.CHAR_FILTER
+                                && 
CHAR_REPLACE_FILTER.equals(sortedProps.get(IndexPolicy.PROP_TYPE))) {
+                            String replacement = sortedProps.getOrDefault(
+                                    PROP_REPLACEMENT, 
CHAR_REPLACE_DEFAULT_REPLACEMENT);
+                            String pattern = canonicalizeCharReplacePattern(
+                                    sortedProps.getOrDefault(PROP_PATTERN, 
CHAR_REPLACE_DEFAULT_PATTERN),
+                                    replacement, fold);
+                            if (pattern.isEmpty()) {
+                                return "";
+                            }
+                            if (isCharReplaceDefault(pattern, replacement, 
fold)) {
+                                // Restating the factory defaults is the bare 
built-in reference.
+                                sortedProps.remove(PROP_PATTERN);
+                                sortedProps.remove(PROP_REPLACEMENT);
+                            } else {
+                                sortedProps.put(PROP_PATTERN, pattern);
+                                sortedProps.put(PROP_REPLACEMENT, replacement);
+                            }
+                        }
+                        if (normalizedType != null && sortedProps.size() == 1) 
{
+                            return normalizedType;
+                        }
+                        return sortedProps.toString();
+                    }
+                }
+            }
+        } catch (RuntimeException e) {
+            // Fall through to built-in resolution or the original name.
+        }
+
+        String normalizedName = normalizeBuiltinComponentName(name, 
expectedType);
+        return "empty".equals(normalizedName) ? "" : normalizedName == null ? 
name : normalizedName;
+    }
+
+    /** Whether this canonical char_replace configuration is what a bare 
built-in reference gets. */
+    private static boolean isCharReplaceDefault(String pattern, String 
replacement, FoldContext fold) {
+        return CHAR_REPLACE_DEFAULT_REPLACEMENT.equals(replacement)
+                && canonicalizeCharReplacePattern(
+                        CHAR_REPLACE_DEFAULT_PATTERN, replacement, 
fold).equals(pattern);
+    }
+
+    private static void canonicalizeEffectiveComponentProperties(
+            TreeMap<String, String> properties, String type, 
IndexPolicyTypeEnum expectedType) {
+        if ("pinyin".equals(type)) {
+            removeBooleanDefaults(properties, true,
+                    "keep_first_letter", "keep_full_pinyin", 
"keep_none_chinese",
+                    "keep_none_chinese_together", 
"keep_none_chinese_in_first_letter",
+                    "lowercase", "trim_whitespace", "ignore_pinyin_offset",
+                    "none_chinese_pinyin_tokenize");
+            removeBooleanDefaults(properties, false,
+                    "keep_separate_first_letter", "keep_joined_full_pinyin", 
"keep_original",
+                    "keep_none_chinese_in_joined_full_pinyin", 
"remove_duplicated_term",
+                    "fixed_pinyin_offset", "keep_separate_chinese");
+            removeIntegerDefault(properties, "limit_first_letter_length", 16);
+            canonicalizePinyinDependencies(properties, expectedType);
+            return;
+        }
+
+        if (expectedType == IndexPolicyTypeEnum.TOKEN_FILTER) {
+            if ("asciifolding".equals(type)) {
+                removeBooleanDefaults(properties, false, "preserve_original");
+            } else if ("word_delimiter".equals(type)) {
+                removeBooleanDefaults(properties, true, "generate_word_parts", 
"generate_number_parts",
+                        "split_on_case_change", "split_on_numerics", 
"stem_english_possessive");
+                removeBooleanDefaults(properties, false, "catenate_words", 
"catenate_numbers",
+                        "catenate_all", "preserve_original");
+                canonicalizeWordSet(properties, "protected_words");
+                canonicalizeTypeTable(properties);
+            } else if ("icu_normalizer".equals(type)) {
+                canonicalizeIcuNormalizerDefaults(properties, false);
+            }
+            return;
+        }
+
+        if (expectedType == IndexPolicyTypeEnum.CHAR_FILTER) {
+            if ("icu_normalizer".equals(type)) {
+                canonicalizeIcuNormalizerDefaults(properties, true);
+            }
+            return;
+        }
+
+        if (expectedType != IndexPolicyTypeEnum.TOKENIZER) {
+            return;
+        }
+        switch (type) {
+            case "ngram":
+            case "edge_ngram":
+                removeIntegerDefault(properties, "min_gram", 1);
+                removeIntegerDefault(properties, "max_gram", 2);
+                canonicalizeWordSet(properties, "token_chars");
+                canonicalizeCustomTokenChars(properties);
+                break;
+            case "standard":
+                removeIntegerDefault(properties, "max_token_length", 255);
+                break;
+            case "char_group":
+                removeIntegerDefault(properties, "max_token_length", 255);
+                canonicalizeTokenizeOnChars(properties);
+                break;
+            case "keyword":
+                // BE only range-checks buffer_size; the emitted term is 
always capped by a constant.
+                properties.remove("buffer_size");
+                break;
+            case "basic":
+                canonicalizeBasicExtraChars(properties);
+                break;
+            default:
+                break;
+        }
+    }
+
+    private static void removeBooleanDefaults(
+            TreeMap<String, String> properties, boolean defaultValue, 
String... keys) {
+        for (String key : keys) {
+            String value = properties.get(key);
+            if (value == null || !("true".equalsIgnoreCase(value) || 
"false".equalsIgnoreCase(value))) {
+                continue;
+            }
+            boolean parsed = Boolean.parseBoolean(value);
+            if (parsed == defaultValue) {
+                properties.remove(key);
+            } else {
+                properties.put(key, Boolean.toString(parsed));
+            }
         }
+    }
 
-        // For custom component, get its properties
+    private static void removeIntegerDefault(
+            TreeMap<String, String> properties, String key, int defaultValue) {
+        String value = properties.get(key);
+        if (value == null) {
+            return;
+        }
         try {
-            Env env = Env.getCurrentEnv();
-            if (env == null || env.getIndexPolicyMgr() == null) {
-                return name;
+            int parsed = Integer.parseInt(value);
+            if (parsed == defaultValue) {
+                properties.remove(key);
+            } else {
+                properties.put(key, Integer.toString(parsed));
             }
+        } catch (NumberFormatException e) {
+            // Invalid policies keep their original identity.
+        }
+    }
 
-            IndexPolicy policy = env.getIndexPolicyMgr().getPolicyByName(name);
-            if (policy == null || policy.getType() != expectedType) {
-                return name;
+    private static void canonicalizeIcuNormalizerDefaults(
+            TreeMap<String, String> properties, boolean hasMode) {
+        String name = properties.get("name");
+        if (name != null) {
+            String normalizedName = name.trim().toLowerCase(Locale.ROOT);
+            if ("nfkc_cf".equals(normalizedName)) {
+                properties.remove("name");
+            } else {
+                properties.put("name", normalizedName);
             }
-            if (policy.isInvalid()) {
-                return "invalid-policy:" + policy.getId() + ":" + 
policy.getName();
+        }
+        String filter = properties.get("unicode_set_filter");
+        if (filter != null && filter.isEmpty()) {
+            // BE treats an explicit empty string like an absent filter.
+            properties.remove("unicode_set_filter");
+        } else if (filter != null) {
+            try {
+                UnicodeSet unicodeSet = new UnicodeSet(filter);
+                if (unicodeSet.isEmpty()) {
+                    properties.remove("unicode_set_filter");
+                } else {
+                    properties.put("unicode_set_filter", 
unicodeSet.toPattern(false));
+                }
+            } catch (IllegalArgumentException e) {
+                // Invalid policies keep their original identity.
+            }
+        }
+        if (hasMode) {
+            canonicalizeIcuNormalizerMode(properties);
+        }
+    }
+
+    private static void canonicalizeIcuNormalizerMode(TreeMap<String, String> 
properties) {
+        removeStringDefault(properties, "mode", "compose");
+        if (!"decompose".equals(properties.get("mode"))) {
+            return;
+        }
+        // BE ignores mode for nfd/nfkd, and nfc/nfkc in decompose mode are 
the same ICU instances.
+        String name = properties.get("name");
+        if ("nfc".equals(name) || "nfd".equals(name)) {
+            properties.put("name", "nfd");
+            properties.remove("mode");
+        } else if ("nfkc".equals(name) || "nfkd".equals(name)) {
+            properties.put("name", "nfkd");
+            properties.remove("mode");
+        }
+    }
+
+    // BE reads these settings as unordered sets of trimmed, non-empty words.
+    private static void canonicalizeWordSet(TreeMap<String, String> 
properties, String key) {
+        String value = properties.get(key);
+        if (value == null) {
+            return;
+        }
+        TreeSet<String> words = new TreeSet<>();
+        for (String word : value.split(",")) {
+            String trimmed = trimAsciiWhitespace(word);
+            if (!trimmed.isEmpty()) {
+                words.add(trimmed);
             }
+        }
+        if (words.isEmpty()) {
+            properties.remove(key);
+        } else {
+            properties.put(key, String.join(",", words));
+        }
+    }
+
+    // BE matches custom token characters as a code point set ORed with the 
named classes.
+    private static void canonicalizeCustomTokenChars(TreeMap<String, String> 
properties) {
+        String value = properties.get("custom_token_chars");
+        if (value == null) {
+            return;
+        }
+        String tokenChars = properties.getOrDefault("token_chars", "");
+        Set<String> classes = new TreeSet<>(List.of(tokenChars.split(",")));
+        StringBuilder canonical = new StringBuilder();
+        value.codePoints().distinct().sorted()
+                .filter(codePoint -> !isCoveredByAsciiClass(codePoint, 
classes))
+                .forEach(canonical::appendCodePoint);
+        if (canonical.length() == 0 && !value.isEmpty() && 
classes.remove("custom")) {
+            properties.remove("custom_token_chars");
+            properties.put("token_chars", String.join(",", classes));
+            return;
+        }
+        properties.put("custom_token_chars", canonical.toString());
+    }
+
+    // BE collects tokenize_on_chars entries into sets and checks the 
categories before the literals.
+    private static void canonicalizeTokenizeOnChars(TreeMap<String, String> 
properties) {
+        List<String> entries = 
parseEntryList(properties.get("tokenize_on_chars"));
+        if (entries == null) {
+            return;
+        }
+        TreeSet<String> canonical = new TreeSet<>(entries);
+        Set<String> classes = new TreeSet<>(canonical);
+        classes.retainAll(CHAR_GROUP_TYPES);
+        canonical.removeIf(entry -> entry.indexOf('\\') < 0
+                && entry.codePointCount(0, entry.length()) == 1
+                && isCoveredByAsciiClass(entry.codePointAt(0), classes));
+        putEntryList(properties, "tokenize_on_chars", canonical);
+    }
 
-            Map<String, String> props = policy.getProperties();
-            if (props == null || props.isEmpty()) {
-                return name;
+    /**
+     * Whether a named character class of the ngram or char_group tokenizer 
already matches this
+     * code point. Only ASCII is judged: its categories never change between 
the ICU versions FE
+     * and BE link against, and the class predicates agree there.
+     */
+    private static boolean isCoveredByAsciiClass(int codePoint, Set<String> 
classes) {
+        if (codePoint >= 128) {
+            return false;
+        }
+        int type = UCharacter.getType(codePoint);
+        for (String name : classes) {
+            switch (name) {
+                case "letter":
+                    if (UCharacter.isLetter(codePoint)) {
+                        return true;
+                    }
+                    break;
+                case "digit":
+                    if (UCharacter.isDigit(codePoint)) {
+                        return true;
+                    }
+                    break;
+                case "whitespace":
+                    if (UCharacter.isWhitespace(codePoint)) {
+                        return true;
+                    }
+                    break;
+                case "punctuation":
+                    if (type == UCharacter.START_PUNCTUATION || type == 
UCharacter.END_PUNCTUATION
+                            || type == UCharacter.OTHER_PUNCTUATION || type == 
UCharacter.CONNECTOR_PUNCTUATION
+                            || type == UCharacter.DASH_PUNCTUATION || type == 
UCharacter.INITIAL_PUNCTUATION
+                            || type == UCharacter.FINAL_PUNCTUATION) {
+                        return true;
+                    }
+                    break;
+                case "symbol":
+                    if (type == UCharacter.CURRENCY_SYMBOL || type == 
UCharacter.MATH_SYMBOL
+                            || type == UCharacter.OTHER_SYMBOL || type == 
UCharacter.MODIFIER_SYMBOL) {
+                        return true;
+                    }
+                    break;
+                default:
+                    break;
             }
+        }
+        return false;
+    }
 
-            // Build identity from sorted properties
-            TreeMap<String, String> sortedProps = new TreeMap<>(props);
-            if (expectedType == IndexPolicyTypeEnum.TOKENIZER
-                    && "ngram".equals(sortedProps.get(IndexPolicy.PROP_TYPE))) 
{
-                // This setting only limits policy creation; it does not 
change emitted tokens.
-                sortedProps.remove(PROP_MAX_NGRAM_DIFF);
+    // BE builds a per-character type map where a later rule for the same 
character wins.
+    private static void canonicalizeTypeTable(TreeMap<String, String> 
properties) {
+        List<String> rules = parseEntryList(properties.get("type_table"));
+        if (rules == null) {
+            return;
+        }
+        TreeMap<Integer, String> types = new TreeMap<>();
+        for (String rule : rules) {
+            int arrow = rule.lastIndexOf("=>");
+            if (arrow < 0 || rule.indexOf('\n') >= 0 || rule.indexOf('\r') >= 
0) {
+                return;
             }
-            return sortedProps.toString();
-        } catch (RuntimeException e) {
-            return name;
+            String character = trimAsciiWhitespace(rule.substring(0, arrow));
+            String type = trimAsciiWhitespace(rule.substring(arrow + 2));
+            // Escaped characters keep the original identity rather than 
reproducing BE unescaping.
+            if (character.indexOf('\\') >= 0 || character.codePointCount(0, 
character.length()) != 1
+                    || !WORD_DELIMITER_TYPES.contains(type)) {
+                return;
+            }
+            types.put(character.codePointAt(0), type);
+        }
+        // An explicit table uses BE's generated classification, including 
when all rules restate it.
+        TreeMap<Integer, String> effectiveTypes = new TreeMap<>(types);
+        effectiveTypes.entrySet().removeIf(
+                entry -> 
entry.getValue().equals(defaultWordDelimiterType(entry.getKey())));
+        if (effectiveTypes.isEmpty() && !types.isEmpty()) {
+            properties.put("type_table", "");
+            return;
+        }
+        List<String> canonicalRules = new ArrayList<>();
+        for (Map.Entry<Integer, String> entry : effectiveTypes.entrySet()) {
+            canonicalRules.add(new String(Character.toChars(entry.getKey())) + 
"=>" + entry.getValue());
+        }
+        putEntryList(properties, "type_table", canonicalRules);
+    }
+
+    /** BE's u_charType classification of an ASCII code point, or null for 
anything else. */
+    private static String defaultWordDelimiterType(int codePoint) {
+        if (codePoint >= 128) {
+            return null;
+        }
+        switch (UCharacter.getType(codePoint)) {
+            case UCharacter.UPPERCASE_LETTER:
+                return "UPPER";
+            case UCharacter.LOWERCASE_LETTER:
+                return "LOWER";
+            case UCharacter.DECIMAL_DIGIT_NUMBER:
+                return "DIGIT";
+            default:
+                return "SUBWORD_DELIM";
+        }
+    }
+
+    /** Parse a bracketed entry list as BE does, or return null for a 
malformed list. */
+    private static List<String> parseEntryList(String value) {
+        if (value == null) {
+            return null;
+        }
+        List<String> entries = new ArrayList<>();
+        String trimmed = trimAsciiWhitespace(value);
+        if (trimmed.isEmpty()) {
+            return entries;
+        }
+        for (String item : ENTRY_SEPARATOR.split(trimmed)) {
+            String entry = trimAsciiWhitespace(item);
+            if (entry.length() < 2 || entry.charAt(0) != '[' || 
entry.charAt(entry.length() - 1) != ']') {
+                return null;
+            }
+            String content = entry.substring(1, entry.length() - 1);
+            if (!content.isEmpty()) {
+                entries.add(content);
+            }
+        }
+        return entries;
+    }
+
+    private static void putEntryList(TreeMap<String, String> properties, 
String key, Collection<String> entries) {
+        if (entries.isEmpty()) {
+            properties.remove(key);
+            return;
+        }
+        StringBuilder canonical = new StringBuilder();
+        for (String entry : entries) {
+            if (canonical.length() > 0) {
+                canonical.append(",");
+            }
+            canonical.append("[").append(entry).append("]");
+        }
+        properties.put(key, canonical.toString());
+    }
+
+    // Trim the same ASCII whitespace that BE trims.
+    private static String trimAsciiWhitespace(String value) {
+        int begin = 0;
+        int end = value.length();
+        while (begin < end && isAsciiWhitespace(value.charAt(begin))) {
+            ++begin;
+        }
+        while (end > begin && isAsciiWhitespace(value.charAt(end - 1))) {
+            --end;
+        }
+        return value.substring(begin, end);
+    }
+
+    private static boolean isAsciiWhitespace(char value) {
+        return value == ' ' || (value >= '\t' && value <= '\r');
+    }
+
+    // BE consumes an ASCII alphanumeric run before it consults extra_chars.
+    private static void canonicalizeBasicExtraChars(TreeMap<String, String> 
properties) {
+        String extraChars = properties.get("extra_chars");
+        if (extraChars == null) {
+            return;
+        }
+        boolean[] present = new boolean[128];
+        for (int i = 0; i < extraChars.length(); ++i) {
+            char value = extraChars.charAt(i);
+            if (value >= present.length) {
+                return;
+            }
+            boolean alphanumeric = (value >= '0' && value <= '9') || (value >= 
'A' && value <= 'Z')
+                    || (value >= 'a' && value <= 'z');
+            present[value] = !alphanumeric;
+        }
+        StringBuilder canonical = new StringBuilder();
+        for (int i = 0; i < present.length; ++i) {
+            if (present[i]) {
+                canonical.append((char) i);
+            }
+        }
+        if (canonical.length() == 0) {
+            properties.remove("extra_chars");
+        } else {
+            properties.put("extra_chars", canonical.toString());
         }
     }
 
+    private static void canonicalizePinyinDependencies(
+            TreeMap<String, String> properties, IndexPolicyTypeEnum 
expectedType) {
+        Boolean keepFirstLetter = effectiveBoolean(properties, 
"keep_first_letter", true);
+        Boolean keepFullPinyin = effectiveBoolean(properties, 
"keep_full_pinyin", true);
+        Boolean keepSeparateFirstLetter = effectiveBoolean(properties, 
"keep_separate_first_letter", false);
+        Boolean keepOriginal = effectiveBoolean(properties, "keep_original", 
false);
+        Boolean keepNoneChinese = effectiveBoolean(properties, 
"keep_none_chinese", true);
+        Boolean keepNoneChineseTogether = effectiveBoolean(properties, 
"keep_none_chinese_together", true);
+        Boolean noneChinesePinyinTokenize = effectiveBoolean(properties, 
"none_chinese_pinyin_tokenize", true);
+        Boolean ignorePinyinOffset = effectiveBoolean(properties, 
"ignore_pinyin_offset", true);
+        Boolean keepJoinedFullPinyin = effectiveBoolean(properties, 
"keep_joined_full_pinyin", false);
+        // Only the pinyin tokenizer also consults 
keep_none_chinese_in_joined_full_pinyin, when no
+        // other setting settles whether it emits an untokenized ASCII buffer.
+        boolean tokenizerReadsJoinedSetting = expectedType == 
IndexPolicyTypeEnum.TOKENIZER
+                && !Boolean.FALSE.equals(keepNoneChinese)
+                && !Boolean.FALSE.equals(keepNoneChineseTogether)
+                && !Boolean.TRUE.equals(noneChinesePinyinTokenize)
+                && !Boolean.TRUE.equals(keepFirstLetter)
+                && !Boolean.TRUE.equals(keepSeparateFirstLetter)
+                && !Boolean.TRUE.equals(keepFullPinyin);
+
+        // The tokenizer only trims its candidates, and without the original 
every candidate is
+        // pinyin or ASCII alphanumerics; the token filter also trims the 
incoming token.
+        if (expectedType == IndexPolicyTypeEnum.TOKENIZER && 
Boolean.FALSE.equals(keepOriginal)) {
+            properties.remove("trim_whitespace");
+        }
+
+        // Without per-character outputs, all remaining candidates have 
position 1.
+        // Deduplicating by term or by term and position has the same effect.
+        if (Boolean.FALSE.equals(keepNoneChinese)
+                && Boolean.FALSE.equals(keepFullPinyin)
+                && Boolean.FALSE.equals(keepSeparateFirstLetter)
+                && Boolean.FALSE.equals(effectiveBoolean(properties, 
"keep_separate_chinese", false))) {
+            properties.remove("remove_duplicated_term");
+        }
+
+        if (Boolean.FALSE.equals(keepFirstLetter)) {
+            properties.remove("limit_first_letter_length");
+            properties.remove("keep_none_chinese_in_first_letter");
+        }
+
+        if (Boolean.FALSE.equals(keepNoneChinese)) {
+            properties.remove("keep_none_chinese_together");
+            properties.remove("none_chinese_pinyin_tokenize");
+        } else if (Boolean.TRUE.equals(keepNoneChinese) && 
Boolean.FALSE.equals(keepNoneChineseTogether)) {
+            // BE emits each ASCII letter on its own here and never tokenizes 
an ASCII buffer.
+            properties.remove("none_chinese_pinyin_tokenize");
+        }
+
+        if (Boolean.TRUE.equals(ignorePinyinOffset)
+                || Boolean.FALSE.equals(keepNoneChinese)
+                || Boolean.FALSE.equals(keepNoneChineseTogether)
+                || Boolean.FALSE.equals(noneChinesePinyinTokenize)) {
+            properties.remove("fixed_pinyin_offset");
+        }
+
+        // The joined full pinyin buffer is only emitted behind 
keep_joined_full_pinyin.
+        if (Boolean.FALSE.equals(keepJoinedFullPinyin) && 
!tokenizerReadsJoinedSetting) {
+            properties.remove("keep_none_chinese_in_joined_full_pinyin");
+        }
+
+        // The ASCII alphabet tokenizer and pinyin dictionary already emit 
lowercase candidates.
+        // The token filter keeps this setting because its fallback can carry 
the source case.
+        if (expectedType == IndexPolicyTypeEnum.TOKENIZER
+                && Boolean.FALSE.equals(keepFirstLetter)
+                && Boolean.FALSE.equals(keepOriginal)
+                && Boolean.FALSE.equals(keepJoinedFullPinyin)

Review Comment:
   Confirmed and fixed in e3a55f591a1.
   
   Going through the candidate sources in `pinyin_tokenizer.cpp`, only five can 
carry source case: the original (`:173-177`), the per-character ASCII output 
(`:124-129`), the untokenized ASCII buffer (`:405`, reached only when 
`none_chinese_pinyin_tokenize` is false, since `walk()` lower-cases 
unconditionally in `pinyin_alphabet_tokenizer.cpp:91-93`), and the first-letter 
and joined aggregates - but those two carry ASCII only when 
`keep_none_chinese_in_first_letter` (`:131-133`) or 
`keep_none_chinese_in_joined_full_pinyin` (`:134-136`) puts it there. Those two 
switches sit outside the `keep_none_chinese` block, so they are independent of 
it. Everything else is dictionary pinyin, which `TONELESS_PINYIN_FORMAT` 
already fixes to lower case.
   
   The condition is derived from all five now, and it is evaluated before the 
canonicalizer removes any of those keys, otherwise the two ASCII switches would 
be read back as their defaults.
   
   You were right that an existing assertion was wrong: `pinyin_joined` versus 
`pinyin_joined_cased` asserted `assertNotEquals`, but that configuration leaves 
`keep_none_chinese_in_joined_full_pinyin` at its default false, so the joined 
aggregate holds dictionary pinyin only and the setting cannot matter. It is 
`assertEquals` now, with `pinyin_joined_ascii` added as the ASCII-enabled 
control and `pinyin_first_letter_dict_only` as the first-letter case. The 
original, untokenized-ASCII and separate-ASCII negatives are unchanged.



##########
fe/fe-core/src/main/java/org/apache/doris/analysis/invertedindex/AnalyzerIdentityBuilder.java:
##########
@@ -226,26 +966,411 @@ private static String resolveTokenFilterIdentity(String 
filterList) {
      * IMPORTANT: Order is preserved because filter order is semantically 
significant.
      */
     private static String resolveCharFilterIdentity(String filterList) {
+        return resolveCharFilterIdentity(filterList, null);
+    }
+
+    private static String resolveCharFilterIdentity(String filterList, 
FoldContext downstreamFold) {
+        ArrayDeque<String> identities = new ArrayDeque<>();
+        walkCharFilters(filterList, downstreamFold, identities);
+        return String.join(",", identities);
+    }
+
+    /**
+     * Resolve the chain from its last filter to its first, collecting 
identities, and return the
+     * case-folding context that a filter placed in front of the chain would 
run in.
+     */
+    private static FoldContext walkCharFilters(
+            String filterList, FoldContext downstreamFold, Deque<String> 
identities) {
+        FoldContext fold = downstreamFold;
         if (Strings.isNullOrEmpty(filterList)) {
-            return "";
+            return fold;
         }
 
-        StringBuilder sb = new StringBuilder();
         String[] filters = filterList.split(",\\s*");
         // DO NOT sort - filter order is semantically significant
 
-        for (int i = 0; i < filters.length; i++) {
-            String filter = filters[i].trim();
-            if (i > 0) {
-                sb.append(",");
+        for (int i = filters.length - 1; i >= 0; --i) {
+            String filterName = filters[i].trim();
+            String filter = resolveComponentIdentity(filterName, 
IndexPolicyTypeEnum.CHAR_FILTER, fold);
+            if (Strings.isNullOrEmpty(filter)) {
+                continue;
             }
+            // Repeating a char_replace filter rewrites the same bytes to the 
same byte again.
+            if (!filter.equals(identities.peekFirst()) || 
!isIdempotentCharFilter(filterName)) {
+                identities.addFirst(filter);
+            }
+            fold = foldContextBefore(filterName, fold);
+        }
+        return fold;
+    }
 
-            if (IndexPolicy.BUILTIN_CHAR_FILTERS.contains(filter)) {
-                sb.append(filter);
-            } else {
-                sb.append(resolveComponentIdentity(filter, 
IndexPolicyTypeEnum.CHAR_FILTER));
+    /**
+     * Context for the filter that runs before this one: a case fold starts a 
fresh context, a
+     * char_replace filter adds the bytes it rewrites, and any other filter 
ends the context.
+     */
+    private static FoldContext foldContextBefore(String filterName, 
FoldContext fold) {
+        FoldContext caseFold = caseFoldingCharFilterContext(filterName);
+        if (caseFold != null) {
+            return caseFold;
+        }
+        if (fold == null) {
+            return null;
+        }
+        boolean[] sourceBytes = charReplaceSourceBytes(filterName);
+        if (sourceBytes == null) {
+            return null;
+        }
+        fold.block(sourceBytes);
+        return fold;
+    }
+
+    /**
+     * Whether the filter is a usable char_replace, which replaces each 
pattern byte with the same
+     * single byte and so leaves the stream unchanged when it runs again.
+     */
+    private static boolean isIdempotentCharFilter(String filterName) {
+        return charReplaceSourceBytes(filterName) != null;
+    }
+
+    /**
+     * Bytes a char_replace filter rewrites, or null for any other filter. A 
bare built-in reference
+     * is instantiated with the factory defaults.
+     */
+    private static boolean[] charReplaceSourceBytes(String filterName) {
+        String pattern = CHAR_REPLACE_DEFAULT_PATTERN;
+        String replacement = CHAR_REPLACE_DEFAULT_REPLACEMENT;
+        IndexPolicy policy = findPolicy(filterName, 
IndexPolicyTypeEnum.CHAR_FILTER);
+        if (policy != null) {
+            if (policy.isInvalid() || policy.getProperties() == null) {
+                return null;
+            }
+            Map<String, String> properties = policy.getProperties();
+            String type = normalizeBuiltinComponentName(
+                    properties.get(IndexPolicy.PROP_TYPE), 
IndexPolicyTypeEnum.CHAR_FILTER);
+            if (!CHAR_REPLACE_FILTER.equals(type)) {
+                return null;
             }
+            pattern = properties.getOrDefault(PROP_PATTERN, 
CHAR_REPLACE_DEFAULT_PATTERN);
+            replacement = properties.getOrDefault(PROP_REPLACEMENT, 
CHAR_REPLACE_DEFAULT_REPLACEMENT);
+        } else if (!CHAR_REPLACE_FILTER.equals(
+                normalizeBuiltinComponentName(filterName, 
IndexPolicyTypeEnum.CHAR_FILTER))) {
+            return null;
         }
-        return sb.toString();
+        // Replacing the single replacement byte with itself leaves the stream 
unchanged.
+        int replacementByte = replacement.length() == 1 && 
replacement.charAt(0) < 128 ? replacement.charAt(0) : -1;
+        boolean[] sourceBytes = new boolean[256];
+        for (int i = 0; i < pattern.length(); ++i) {
+            char patternByte = pattern.charAt(i);
+            if (patternByte < sourceBytes.length && patternByte != 
replacementByte) {
+                sourceBytes[patternByte] = true;
+            }
+        }
+        return sourceBytes;
+    }
+
+    /** The named policy when one exists with the expected type, or null. */
+    private static IndexPolicy findPolicy(String name, IndexPolicyTypeEnum 
expectedType) {
+        if (Strings.isNullOrEmpty(name)) {
+            return null;
+        }
+        try {
+            Env env = Env.getCurrentEnv();
+            if (env != null && env.getIndexPolicyMgr() != null) {
+                IndexPolicy policy = 
env.getIndexPolicyMgr().getPolicyByName(name);
+                if (policy != null && policy.getType() == expectedType) {
+                    return policy;
+                }
+            }
+        } catch (RuntimeException e) {
+            // Treat lookup failures as an unknown policy.
+        }
+        return null;
+    }
+
+    /** Fold context started by a named or built-in case-folding char filter, 
or null for any other filter. */
+    private static FoldContext caseFoldingCharFilterContext(String name) {
+        if (Strings.isNullOrEmpty(name)) {
+            return null;
+        }
+
+        try {
+            Env env = Env.getCurrentEnv();
+            if (env != null && env.getIndexPolicyMgr() != null) {
+                IndexPolicy policy = 
env.getIndexPolicyMgr().getPolicyByName(name);
+                if (policy != null && policy.getType() == 
IndexPolicyTypeEnum.CHAR_FILTER) {
+                    if (policy.isInvalid()) {
+                        return null;
+                    }
+                    Map<String, String> properties = policy.getProperties();
+                    if (properties != null && !properties.isEmpty()) {
+                        String type = normalizeBuiltinComponentName(
+                                properties.get(IndexPolicy.PROP_TYPE), 
IndexPolicyTypeEnum.CHAR_FILTER);
+                        return "icu_normalizer".equals(type) ? 
icuNormalizerFoldContext(properties) : null;
+                    }
+                }
+            }
+        } catch (RuntimeException e) {
+            // Fall through to built-in resolution.
+        }
+
+        return "icu_normalizer".equals(normalizeBuiltinComponentName(name, 
IndexPolicyTypeEnum.CHAR_FILTER))
+                ? FoldContext.unfiltered() : null;
+    }
+
+    /**
+     * Fold context of an icu_normalizer component: the default nfkc_cf form 
folds case over every
+     * code point, or only inside a parsable non-empty unicode_set_filter. 
Null for other forms.
+     */
+    private static FoldContext icuNormalizerFoldContext(Map<String, String> 
properties) {
+        if (!"nfkc_cf".equals(icuNormalizerName(properties))) {
+            return null;
+        }
+        String filter = properties.get("unicode_set_filter");
+        if (filter == null || filter.isEmpty()) {
+            return FoldContext.unfiltered();
+        }
+        try {
+            UnicodeSet unicodeSet = new UnicodeSet(filter);
+            return unicodeSet.isEmpty() ? FoldContext.unfiltered() : new 
FoldContext(unicodeSet.freeze());
+        } catch (IllegalArgumentException e) {
+            return null;
+        }
+    }
+
+    /** Whether an icu_normalizer component leaves ASCII letters as they are. 
*/
+    private static boolean isAsciiCaseTransparentIcuNormalizer(Map<String, 
String> properties) {
+        String name = icuNormalizerName(properties);
+        return "nfc".equals(name) || "nfd".equals(name) || "nfkc".equals(name) 
|| "nfkd".equals(name);
+    }
+
+    private static String icuNormalizerName(Map<String, String> properties) {
+        return properties.getOrDefault("name", 
"nfkc_cf").trim().toLowerCase(Locale.ROOT);
+    }
+
+    /** The outer char filter runs before everything else, so it takes the 
analyzer's fold context. */
+    private static String appendOuterCharFilterIdentity(
+            String analyzerIdentity, Map<String, String> properties, 
FoldContext fold) {
+        String type = 
properties.get(InvertedIndexProperties.INVERTED_INDEX_PARSER_CHAR_FILTER_TYPE);
+        String pattern = 
properties.get(InvertedIndexProperties.INVERTED_INDEX_PARSER_CHAR_FILTER_PATTERN);
+        if (!"char_replace".equals(type) || Strings.isNullOrEmpty(pattern)) {
+            return analyzerIdentity;
+        }
+        String replacement = properties.getOrDefault(
+                
InvertedIndexProperties.INVERTED_INDEX_PARSER_CHAR_FILTER_REPLACEMENT, " ");
+        String canonicalPattern = canonicalizeCharReplacePattern(pattern, 
replacement, fold);
+        if (canonicalPattern.isEmpty()) {
+            return analyzerIdentity;
+        }
+        return analyzerIdentity + "|outer_char_filter=char_replace:"
+                + canonicalPattern.length() + ":" + canonicalPattern + ":"
+                + replacement.length() + ":" + replacement + ";";
+    }
+
+    /**
+     * Canonicalize the ASCII pattern to the BE filter's byte set.
+     * Order, duplicate bytes, and replacements of a byte with itself do not 
change the stream.
+     */
+    private static String canonicalizeCharReplacePattern(
+            String pattern, String replacement, FoldContext fold) {
+        if (replacement.length() != 1) {
+            return pattern;
+        }
+        char replacementByte = replacement.charAt(0);
+        boolean[] replacedBytes = new boolean[256];
+        for (int i = 0; i < pattern.length(); ++i) {
+            char patternByte = pattern.charAt(i);
+            if (patternByte < replacedBytes.length && patternByte != 
replacementByte) {
+                replacedBytes[patternByte] = true;
+            }
+        }
+        if (fold != null && replacementByte >= 'a' && replacementByte <= 'z') {
+            // The downstream fold maps the upper-case byte to the replacement 
anyway.
+            int upperByte = replacementByte - ('a' - 'A');
+            if (fold.foldsByte(upperByte, replacementByte)) {
+                replacedBytes[upperByte] = false;
+            }
+        } else if (fold != null && replacementByte >= 'A' && replacementByte 
<= 'Z') {
+            // The downstream fold maps the replacement back to the lower-case 
byte it replaced.
+            int lowerByte = replacementByte + ('a' - 'A');
+            if (fold.foldsByte(replacementByte, lowerByte)) {
+                replacedBytes[lowerByte] = false;
+            }
+        }
+
+        StringBuilder canonical = new StringBuilder();
+        for (int i = 0; i < replacedBytes.length; ++i) {
+            if (replacedBytes[i]) {
+                canonical.append((char) i);
+            }
+        }
+        return canonical.toString();
+    }
+
+    private static FoldContext builtinIkFoldContext(String analyzerIdentity) {
+        return isDefaultLowercaseBuiltinIkIdentity(analyzerIdentity) ? 
FoldContext.unfiltered() : null;

Review Comment:
   Confirmed and fixed in e3a55f591a1.
   
   `IKAnalyzer` builds `Configuration(true, true)` (`IKAnalyzer.h:26-30`), and 
`setEnableLowercase` has no caller anywhere in the tree, so 
`config_->isEnableLowercase()` is always true. The later `set_lowercase(false)` 
only writes the base class field that `IKTokenizer` consults for its own 
regularization. The fold that matters happens earlier: `AnalyzeContext` decodes 
with that always-true flag (`AnalyzeContext.cpp:70-72,83-86`) and 
`CharacterUtil.cpp:93-95` rewrites `A-Z` in place inside `segment_buff_`, which 
is where the lexeme text is copied from (`AnalyzeContext.cpp:244-245`).
   
   So the fold context is unconditional now, and the assertion that assumed 
otherwise - plain IK with `lower_case=false` versus the same analyzer with an 
outer `A -> a` - has been flipped to `assertEquals`. A new negative, an outer 
rewrite of `-` to a space under the same `lower_case=false`, keeps a char 
filter that does change the terms distinct.
   
   One boundary worth noting: full-width upper case is converted to half width 
and only then lower-cased according to `use_lowercase` 
(`CharacterUtil.cpp:110-119`), so it is not folded when `lower_case=false`. The 
char_replace model is single-byte and cannot express that rewrite, so it does 
not affect this identity.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to