[ 
https://issues.apache.org/jira/browse/FLINK-40871?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Ramin Gharib resolved FLINK-40871.
----------------------------------
    Resolution: Fixed

merged as: https://github.com/apache/flink/pull/29370

> Avoid the UTF-8 round trip when casting strings and decimals to VARIANT
> -----------------------------------------------------------------------
>
>                 Key: FLINK-40871
>                 URL: https://issues.apache.org/jira/browse/FLINK-40871
>             Project: Flink
>          Issue Type: Sub-task
>          Components: Table SQL / Runtime
>            Reporter: Ramin Gharib
>            Assignee: Ramin Gharib
>            Priority: Major
>              Labels: pull-request-available
>
> FLINK-40825 added \{{CAST}} from primitive types to VARIANT. Two of its 
> runtime helpers do more work per record than needed.
> h3. Problem
> *Strings.* \{{VariantCastUtils#fromString}} calls \{{StringData#toString()}}, 
> which decodes the UTF-8 bytes into a Java String. 
> \{{BinaryVariantInternalBuilder#appendString}} then encodes the String back 
> to UTF-8, and \{{build()}} copies the result once more. The bytes that land 
> in the VARIANT are, for valid input, the bytes we started with.
> *Decimals.* \{{VariantCastUtils#fromDecimal}} calls 
> \{{DecimalData#toBigDecimal()}}. A compact decimal, with a precision of 18 or 
> less, already holds an unscaled long. The BigDecimal is only built for 
> \{{appendDecimal}} to take it apart again.
> h3. The catch: UTF-8 validity
> The Variant spec requires string values to be valid UTF-8. A 
> \{{BinaryStringData}} is not guaranteed to hold valid UTF-8. Today the decode 
> step hides this: \{{StringUtf8Utils#decodeUTF8}} falls back to \{{new 
> String(bytes, UTF_8)}}, which replaces every malformed sequence with U+FFFD. 
> So a VARIANT string is always valid UTF-8, if lossy.
> Copying \{{BinaryStringData#toBytes()}} straight into the builder would put 
> invalid bytes into the VARIANT. Any byte path therefore needs a validation 
> step.
> h3. Proposal
> * Add an \{{appendString(byte[] utf8)}} overload to 
> \{{BinaryVariantInternalBuilder}}, and let \{{appendString(String)}} delegate 
> to it.
> * In \{{fromString}}, validate the bytes in one pass. Valid bytes go to the 
> new overload. Invalid bytes fall back to today's \{{toString()}} path. This 
> keeps the result identical to today for every input, and a scan is cheaper 
> than a decode plus an encode.
> * Add a compact path for decimals that writes the unscaled long directly, 
> picking decimal4 or decimal8 the same way \{{appendDecimal}} does today.
> Rejecting invalid UTF-8 instead of replacing it is an option too. It would 
> change behavior and make every string cast to VARIANT fallible, so it should 
> be a separate decision.
> h3. Acceptance
> * For valid UTF-8, the stored VARIANT is byte-for-byte identical to today's.
> * Invalid UTF-8 still stores U+FFFD, exactly as today.
> * Compact and non-compact decimals produce the same VARIANT as today, at 
> every width boundary.
> * A micro benchmark shows the gain for both casts.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to