Zoltán Borók-Nagy created IMPALA-15423:
------------------------------------------

             Summary: Non-ASCII strings are truncated of Iceberg metadata tables
                 Key: IMPALA-15423
                 URL: https://issues.apache.org/jira/browse/IMPALA-15423
             Project: IMPALA
          Issue Type: Bug
          Components: Backend
            Reporter: Zoltán Borók-Nagy
            Assignee: Zoltán Borók-Nagy


There are two problems in the string/binary path of IcebergRowReader
(be/src/exec/iceberg-metadata/iceberg-row-reader.cc).

h3. 1. Non-ASCII strings are truncated

IcebergRowReader::WriteStringOrBinarySlot() copies {{jbuffer_guard.get_size()}} 
bytes. For
strings, JniBufferGuard<jstring>::create() in be/src/util/jni-util.cc computes 
the size
and the buffer from two different encodings:

{code:cpp}
size = env->GetStringLength(jbuffer);                // UTF-16 code units
buffer = env->GetStringUTFChars(jbuffer, &is_copy);  // modified UTF-8 bytes
{code}

Every non-ASCII character takes more UTF-8 bytes than UTF-16 units, so the 
copied value
is cut short. For example, 'Zürich' (6 units, 7 bytes) comes back as 'Züric', 
and CJK
text loses 2 bytes per character. Also, modified UTF-8 is not standard UTF-8:
supplementary characters are encoded as 6-byte surrogate pairs and NUL as C0 
80. Using
the right length alone would therefore still return non-standard bytes for 
emoji and
similar characters.

This affects every STRING column and map value read through the metadata 
scanner, e.g.
partition values in {{`partitions`}} and {{`files`}}, and lower/upper bounds in
readable_metrics.

{code:sql}
CREATE TABLE ice_utf8 (city STRING) PARTITIONED BY SPEC (city) STORED AS 
ICEBERG;
INSERT INTO ice_utf8 VALUES ('Zürich');
SELECT `partition` FROM ice_utf8.`partitions`;   -- expected {"city":"Zürich"}
{code}

The other JniUtfCharGuard users (fe-support.cc, logging-support.cc) only use 
get() as a
NUL-terminated string, so as far as I can see this reader is the only caller 
hit by the
size mismatch.

h3. 2. JNI local references leak per value

* WriteStringOrBinarySlot() never deletes {{jbuffer}}: the jstring returned by
toString(), or the byte[] from ConvertJavaByteBufferToByteArray().
* WriteMapKeyAndValue() never deletes the {{key}} and {{value}} objects 
returned by
GetNextMapKeyAndValue().

The fragment thread is a native thread attached to the JVM and pushes no local 
frame.
These references therefore live until the thread detaches, keeping one Java 
object
reachable for every string/binary value and map entry of the whole scan. On
{{`files`}}, {{`entries`}} and {{`all_*`}} tables of large tables, this pins a 
lot of heap
in the coordinator's JVM.

h3. Proposed fix

* Produce the bytes on the Java side with 
{{String.getBytes(StandardCharsets.UTF_8)}} and
read them through JniByteArrayGuard. Alternatively, fix JniBufferGuard<jstring> 
to use
GetStringUTFLength and document the modified-UTF-8 caveat.
* DeleteLocalRef {{jbuffer}}, {{key}} and {{value}} after use, or push/pop a 
JniLocalFrame
per row in IcebergMetadataScanNode::GetNext().
* Add EE tests with non-ASCII partition values, including a supplementary 
character.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to