leaves12138 commented on code in PR #958:
URL: https://github.com/apache/paimon-rust/pull/958#discussion_r4109574165


##########
crates/paimon/src/spec/partition_utils.rs:
##########
@@ -316,6 +326,61 @@ fn format_partition_value(
     Ok(value)
 }
 
+/// Decode as Java's UTF-8 decoder does for `BinaryString.toString()`.
+/// Rust's lossy decoder replaces each byte of a UTF-8 encoded surrogate,
+/// while Java replaces the complete malformed surrogate sequence once.
+fn decode_java_utf8(mut bytes: &[u8]) -> String {
+    let mut decoded = String::with_capacity(bytes.len());
+    loop {
+        match std::str::from_utf8(bytes) {
+            Ok(valid) => {
+                decoded.push_str(valid);
+                break;
+            }
+            Err(error) => {
+                let valid_len = error.valid_up_to();
+                
decoded.push_str(std::str::from_utf8(&bytes[..valid_len]).unwrap());
+                bytes = &bytes[valid_len..];
+
+                let malformed_len =
+                    if bytes.len() >= 2 && bytes[0] == 0xED && 
(0xA0..=0xBF).contains(&bytes[1]) {
+                        // The first two bytes denote a surrogate. Java 
consumes
+                        // its third byte too when it is a continuation byte.
+                        if bytes.get(2).is_some_and(|byte| byte & 0xC0 == 
0x80) {
+                            3
+                        } else {
+                            2
+                        }
+                    } else {
+                        error.error_len().unwrap_or(bytes.len())
+                    };
+                decoded.push('\u{FFFD}');
+                bytes = &bytes[malformed_len..];
+            }
+        }
+    }
+    decoded
+}
+
+/// Java `StringUtils.isNullOrWhitespaceOnly` checks each UTF-16 code unit with
+/// `Character.isWhitespace`; its whitespace set differs from Rust `str::trim`.
+fn is_java_whitespace_only(value: &str) -> bool {
+    value.chars().all(|ch| {
+        matches!(
+            ch,
+            '\u{0009}'..='\u{000D}'
+                | '\u{001C}'..='\u{0020}'
+                | '\u{1680}'
+                | '\u{180E}'

Review Comment:
   [P2] Handle the JDK-dependent U+180E partition case explicitly
   
   Including U+180E hardcodes the JDK 8 whitespace definition, but Java 11/17 
`Character.isWhitespace('\u180e')` is false. For a non-legacy VARBINARY 
partition containing bytes `E1 A0 8E`, this code now writes under 
`bin=__DEFAULT_PARTITION__/`, while the actual Paimon 
`InternalRowPartitionComputer` running on JDK 11 or 17 generates 
`bin=<U+180E>/`. I reproduced the missing Java-derived file path after real 
Rust writes and commits for both append and primary-key tables. The repository 
has JDK 11/17 lanes, so matching only the JDK 8 predicate is not sufficient to 
claim general Java path compatibility. This is an inherent Java-version 
difference, not another UTF-8 decoding error: please define and enforce a 
compatible runtime policy, provide an explicit compatibility setting, or reject 
the ambiguous whitespace-only partition value before committing it. Simply 
removing U+180E would reverse which JDK is incompatible, so the compatibility 
boundary should be deliberate rather than silently
  accepted.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to