Zoltán Borók-Nagy created IMPALA-15423:
------------------------------------------
Summary: Non-ASCII strings are truncated of Iceberg metadata tables
Key: IMPALA-15423
URL: https://issues.apache.org/jira/browse/IMPALA-15423
Project: IMPALA
Issue Type: Bug
Components: Backend
Reporter: Zoltán Borók-Nagy
Assignee: Zoltán Borók-Nagy
There are two problems in the string/binary path of IcebergRowReader
(be/src/exec/iceberg-metadata/iceberg-row-reader.cc).
h3. 1. Non-ASCII strings are truncated
IcebergRowReader::WriteStringOrBinarySlot() copies {{jbuffer_guard.get_size()}}
bytes. For
strings, JniBufferGuard<jstring>::create() in be/src/util/jni-util.cc computes
the size
and the buffer from two different encodings:
{code:cpp}
size = env->GetStringLength(jbuffer); // UTF-16 code units
buffer = env->GetStringUTFChars(jbuffer, &is_copy); // modified UTF-8 bytes
{code}
Every non-ASCII character takes more UTF-8 bytes than UTF-16 units, so the
copied value
is cut short. For example, 'Zürich' (6 units, 7 bytes) comes back as 'Züric',
and CJK
text loses 2 bytes per character. Also, modified UTF-8 is not standard UTF-8:
supplementary characters are encoded as 6-byte surrogate pairs and NUL as C0
80. Using
the right length alone would therefore still return non-standard bytes for
emoji and
similar characters.
This affects every STRING column and map value read through the metadata
scanner, e.g.
partition values in {{`partitions`}} and {{`files`}}, and lower/upper bounds in
readable_metrics.
{code:sql}
CREATE TABLE ice_utf8 (city STRING) PARTITIONED BY SPEC (city) STORED AS
ICEBERG;
INSERT INTO ice_utf8 VALUES ('Zürich');
SELECT `partition` FROM ice_utf8.`partitions`; -- expected {"city":"Zürich"}
{code}
The other JniUtfCharGuard users (fe-support.cc, logging-support.cc) only use
get() as a
NUL-terminated string, so as far as I can see this reader is the only caller
hit by the
size mismatch.
h3. 2. JNI local references leak per value
* WriteStringOrBinarySlot() never deletes {{jbuffer}}: the jstring returned by
toString(), or the byte[] from ConvertJavaByteBufferToByteArray().
* WriteMapKeyAndValue() never deletes the {{key}} and {{value}} objects
returned by
GetNextMapKeyAndValue().
The fragment thread is a native thread attached to the JVM and pushes no local
frame.
These references therefore live until the thread detaches, keeping one Java
object
reachable for every string/binary value and map entry of the whole scan. On
{{`files`}}, {{`entries`}} and {{`all_*`}} tables of large tables, this pins a
lot of heap
in the coordinator's JVM.
h3. Proposed fix
* Produce the bytes on the Java side with
{{String.getBytes(StandardCharsets.UTF_8)}} and
read them through JniByteArrayGuard. Alternatively, fix JniBufferGuard<jstring>
to use
GetStringUTFLength and document the modified-UTF-8 caveat.
* DeleteLocalRef {{jbuffer}}, {{key}} and {{value}} after use, or push/pop a
JniLocalFrame
per row in IcebergMetadataScanNode::GetNext().
* Add EE tests with non-ASCII partition values, including a supplementary
character.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]