leaves12138 commented on code in PR #958:
URL: https://github.com/apache/paimon-rust/pull/958#discussion_r4109574165
##########
crates/paimon/src/spec/partition_utils.rs:
##########
@@ -316,6 +326,61 @@ fn format_partition_value(
Ok(value)
}
+/// Decode as Java's UTF-8 decoder does for `BinaryString.toString()`.
+/// Rust's lossy decoder replaces each byte of a UTF-8 encoded surrogate,
+/// while Java replaces the complete malformed surrogate sequence once.
+fn decode_java_utf8(mut bytes: &[u8]) -> String {
+ let mut decoded = String::with_capacity(bytes.len());
+ loop {
+ match std::str::from_utf8(bytes) {
+ Ok(valid) => {
+ decoded.push_str(valid);
+ break;
+ }
+ Err(error) => {
+ let valid_len = error.valid_up_to();
+
decoded.push_str(std::str::from_utf8(&bytes[..valid_len]).unwrap());
+ bytes = &bytes[valid_len..];
+
+ let malformed_len =
+ if bytes.len() >= 2 && bytes[0] == 0xED &&
(0xA0..=0xBF).contains(&bytes[1]) {
+ // The first two bytes denote a surrogate. Java
consumes
+ // its third byte too when it is a continuation byte.
+ if bytes.get(2).is_some_and(|byte| byte & 0xC0 ==
0x80) {
+ 3
+ } else {
+ 2
+ }
+ } else {
+ error.error_len().unwrap_or(bytes.len())
+ };
+ decoded.push('\u{FFFD}');
+ bytes = &bytes[malformed_len..];
+ }
+ }
+ }
+ decoded
+}
+
+/// Java `StringUtils.isNullOrWhitespaceOnly` checks each UTF-16 code unit with
+/// `Character.isWhitespace`; its whitespace set differs from Rust `str::trim`.
+fn is_java_whitespace_only(value: &str) -> bool {
+ value.chars().all(|ch| {
+ matches!(
+ ch,
+ '\u{0009}'..='\u{000D}'
+ | '\u{001C}'..='\u{0020}'
+ | '\u{1680}'
+ | '\u{180E}'
Review Comment:
[P2] Handle the JDK-dependent U+180E partition case explicitly
Including U+180E hardcodes the JDK 8 whitespace definition, but Java 11/17
`Character.isWhitespace('\u180e')` is false. For a non-legacy VARBINARY
partition containing bytes `E1 A0 8E`, this code now writes under
`bin=__DEFAULT_PARTITION__/`, while the actual Paimon
`InternalRowPartitionComputer` running on JDK 11 or 17 generates
`bin=<U+180E>/`. I reproduced the missing Java-derived file path after real
Rust writes and commits for both append and primary-key tables. The repository
has JDK 11/17 lanes, so matching only the JDK 8 predicate is not sufficient to
claim general Java path compatibility. This is an inherent Java-version
difference, not another UTF-8 decoding error: please define and enforce a
compatible runtime policy, provide an explicit compatibility setting, or reject
the ambiguous whitespace-only partition value before committing it. Simply
removing U+180E would reverse which JDK is incompatible, so the compatibility
boundary should be deliberate rather than silently
accepted.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]