尹茂椿萱 created COLLECTIONS-899:
--------------------------------
Summary: CaseInsensitiveMap incorrectly converts U+03D0 (ϐ) to
U+03B2 (β), causing distinct keys to collide
Key: COLLECTIONS-899
URL: https://issues.apache.org/jira/browse/COLLECTIONS-899
Project: Commons Collections
Issue Type: Bug
Components: Map
Affects Versions: 4.6.0
Environment: Apache Commons Collections 4.6.0
Java: 17
OS: Windows 11
Reporter: 尹茂椿萱
{{CaseInsensitiveMap}} performs an incorrect Unicode lowercase conversion for
some characters.
The class-level Javadoc states that keys are converted to all lowercase in a
locale-independent fashion using information from the Unicode data file.
However, the current implementation converts each character using:
{{Character.toLowerCase(Character.toUpperCase(chars[i]))}}
This is not equivalent to Unicode lowercase conversion.
For example, the Unicode character U+03D0 (GREEK BETA SYMBOL, {{{}ϐ{}}}) is
unchanged by lowercase conversion:
{{"\u03D0".toLowerCase(Locale.ROOT)}}
returns {{{}ϐ{}}}.
However, the current implementation converts it as follows:
{{ϐ (U+03D0)
-> Character.toUpperCase()
β (U+03B2)
-> Character.toLowerCase()
β (U+03B2)}}
As a result, {{ϐ}} and {{β}} are incorrectly normalized to the same key.
This causes observable incorrect behavior in the public {{CaseInsensitiveMap}}
API. For example:
{{CaseInsensitiveMap<String, String> map = new CaseInsensitiveMap<>();
map.put("\u03D0", "beta-symbol");
map.put("\u03B2", "greek-beta");
assertEquals(2, map.size());}}
The expected size is {{{}2{}}}, because the two input keys are distinct and are
not lowercase-equivalent according to Unicode lowercase mapping.
With the current implementation, the actual size is {{1}} because both keys are
converted to {{{}β{}}}.
The previous implementation used:
{{key.toString().toLowerCase()}}
and preserved {{ϐ}} as {{{}ϐ{}}}, so this collision did not occur.
This appears to be a regression introduced by the character-wise case
conversion implementation.
*Expected behavior*
{{CaseInsensitiveMap}} should convert keys according to locale-independent
Unicode lowercase semantics. In particular:
{{U+03D0 (ϐ) -> U+03D0 (ϐ)
U+03B2 (β) -> U+03B2 (β)}}
Therefore, inserting both keys should retain two distinct entries.
*Actual behavior*
The current implementation converts:
{{U+03D0 (ϐ) -> U+03B2 (β)
U+03B2 (β) -> U+03B2 (β)}}
Consequently, inserting both keys results in only one map entry.
*Steps to reproduce*
# Create a {{{}CaseInsensitiveMap{}}}.
# Insert {{"\u03D0"}} as a key.
# Insert {{"\u03B2"}} as a key.
# Check the map size.
Reproducer:
{{CaseInsensitiveMap<String, String> map = new CaseInsensitiveMap<>();
map.put("\u03D0", "beta-symbol");
map.put("\u03B2", "greek-beta");
System.out.println(map.size());}}
Actual result:
{{1}}
Expected result:
{{2}}
--
This message was sent by Atlassian Jira
(v8.20.10#820010)