[
https://issues.apache.org/jira/browse/FLINK-40218?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Ramin Gharib updated FLINK-40218:
---------------------------------
Description:
The only way to build a \{{Variant}} from JSON today is from a \{{String}}.
\{{PARSE_JSON}} and the internal
\{{BinaryVariantInternalBuilder.parseJson(String)}} both create a Jackson
parser over the text and walk its tokens in \{{buildJson}}.
A caller that already holds parsed JSON cannot use that work. It must write its
data back to a \{{String}}, which flink-core then parses a second time:
{code}
bytes --(caller parses)--> tree --(toString)--> String --(flink-core parses
again)--> Variant
{code}
Two things keep callers on this path:
# *There is no entry point that takes a parser.* \{{parseJson(JsonParser,
boolean)}} exists but is private.
# *Jackson types do not cross shading boundaries.* The builder uses Flink's
shaded \{{org.apache.flink.shaded.jackson2...JsonParser}}. A caller with its
own, unshaded Jackson cannot pass in its parser or tree. A \{{String}} is the
only type both sides share.
*Proposed change*
Two entry points on the \{{@Internal}} \{{BinaryVariantInternalBuilder}}:
||The caller holds||Entry point||
|Flink's shaded Jackson \{{JsonParser}}|\{{parseJson(JsonParser, boolean)}},
now public|
|JSON in another form, such as a tree from its own Jackson|walk it with the
existing \{{append*}}, \{{addKey}} and \{{finishWritingObject}} methods, and
call the new \{{appendJsonNumber(String)}} for numbers|
{\{parseJson(JsonParser, boolean)}} expects the parser on the value's first
token and leaves it on the value's last token, so a format can read a VARIANT
field in the middle of a record.
{\{appendJsonNumber(String)}} reads the literal with the same Jackson factory
and number code as \{{PARSE_JSON}}. A number is therefore stored byte for byte
like \{{PARSE_JSON}} stores it: the smallest integer, else a decimal, else a
finite double. Anything that is not exactly one JSON number fails.
A walker over an unshaded Jackson tree then looks li
{code:java}
case NUMBER:
if (node.isIntegralNumber() && node.canConvertToLong()) {
builder.appendNumeric(node.longValue());
} else {
builder.appendJsonNumber(node.asText());
}
\{code}
*Design notes*
* *Push instead of a token-source interface.* An earlier proposal added a
pull-based \{{VariantJsonSource}} interface that builder reads from. It is
replaced by one push methocers push into the builder too, such
as\{{ToVariantConverter}} for \{{CAST(... AS VARIANT)}}. One method is a
smaller surface than an interface with an enum. {{PARSkeeps its own code path
unchanged.
* *A String literal, not a \{{Number}}.* \{{PARSE_JSON}} picks the type from
how a number is written, for example an exponenDOUBLE. A \{{BigDecimal}} cannot
tell \{{1e-7}} from \{E_JSON}} stores as DOUBLE and DECIMAL. The text keeps
that information.
* *No hand-written number grammar.*
\{{appendJsonNume the literal, so it rejects {{+5}}, \{{0x1p4}},\{{1.5f}} and
\{{1,5}} exactly like \{{PARSE_JSON}} does.
*Scope*
Moving the \{{json}} format onto the new parser entry point is split into
FLINK-XXXXX. That change is not behavior-neutral: format rounds VARIANT floats
through a \{{double}} to0}} becomes the string \{{"Infinity"}}. It needs itsown
release note.
*Compatibility*
Additive and \{{@Internal}}. \{{PARSE_JSON}}, \{{TRY_PARSE_JSON}} and the
\{{json}} format do not change. \{{addKey}} now looks once instead of twice.
*Verifying this change*
* \{{appendJsonNumber}} is byte-for-byte equal to \{{parseJson(String)}} for
integers at every width, values beyond a long, decimals, exponents and
surrounding whitespace.
* It rejects non-JSON literals, other JSON values such as \{{"5"}} or
\{{[1]}}, trailing content such as \{{5 6}}, and numbers outside the double
range.
* A parser read in the middle of a document is left on the value's last token.
was:
*Description*
The only way to build a {{Variant}} from JSON today is from a {{{}String{}}}.
{{PARSE_JSON}} and the internal
{{BinaryVariantInternalBuilder.parseJson(String)}} both take a string, create a
Jackson {{JsonParser}} over it, and walk the tokens in {{{}buildJson{}}}.
A consumer that already holds parsed JSON cannot use that work. It must
serialize its data back to a {{String}} and hand it to Flink, which parses it a
second time. This shows up wherever a format or connector wants a VARIANT
column: the format has already parsed the record into a tree or a parser of its
own, yet it is forced to {{toString()}} and let flink-core re-parse.
Two things drive this:
# *A redundant pass.* The value is parsed by the caller, serialized to a
{{{}String{}}}, then parsed again by flink-core.
# *The Jackson type is not portable.* {{buildJson}} is written against Flink's
shaded {{{}org.apache.flink.shaded.jackson2...JsonParser{}}}. A caller that
uses a differently relocated or unshaded Jackson cannot pass its own parser in.
A {{String}} is the only type that crosses that boundary, which is why the
round-trip exists.
Concretely, for one VARIANT value read through a JSON-based format:
{code:java}
Current:
bytes ──(caller parses)──▶ tree ──(toString)──▶ String ──(flink re-parses)──▶
tokens ──▶ Variant
[needed] [wasted] [wasted] {code}
*Proposed change*
Introduce a small token-source abstraction in
{{org.apache.flink.types.variant}} and drive {{buildJson}} off it. The concrete
Jackson parser becomes one implementation of that abstraction rather than a
hard dependency.
{code:java}
// flink-core: org.apache.flink.types.variant
@Internal
public interface VariantJsonSource {
/** Advance to the next token and return it. Returns END_INPUT once input
is exhausted. */
Token next() throws IOException;
/** Field name of the current FIELD_NAME token. */
String fieldName() throws IOException;
/** Text of the current STRING token. */
String stringValue() throws IOException;
/**
* Raw literal of the current NUMBER token, e.g. "1e5", "100000", "3.14".
* The builder decides long vs decimal vs double from this text, so numeric
* classification stays in one place and every source behaves identically.
*/
String numberText() throws IOException;
enum Token {
START_OBJECT, END_OBJECT,
START_ARRAY, END_ARRAY,
FIELD_NAME,
STRING, NUMBER, TRUE, FALSE, NULL,
END_INPUT
}
} {code}
{code:java}
// BinaryVariantInternalBuilder
public static BinaryVariant parseJson(VariantJsonSource source, boolean
allowDuplicateKeys) throws IOException {
// buildJson, expressed purely against VariantJsonSource
}
{code}
With the interface a JSON-based format wraps the value it already parsed and
builds the Variant in a single walk:
{code:java}
Proposed:
bytes ──(caller parses)──▶ tree ──(wrap as source)──▶ tokens ──▶ Variant
[needed] [no serialize] [no re-parse] {code}
*Design notes*
* Number classification stays in the builder. The source exposes the raw
number literal through {{{}numberText(){}}}, so the long-vs-decimal-vs-double
decision, including the rule that scientific notation goes to {{{}double{}}},
lives in one place. Sources do not re-implement it and cannot diverge.
* Boolean values are carried by the {{TRUE}} / {{FALSE}} tokens, so no
separate value accessor is needed.
* The interface uses only JDK types plus its own enum. It names no Jackson
type, shaded or otherwise, so any caller can implement it regardless of how its
Jackson is relocated.
* The exact signature is open. {{numberText}} versus typed number accessors is
the main choice, and the text form is preferred because it is the only shape
that preserves the current numeric routing without leaking policy into each
source. This can be finalized in the PR.
*Compatibility*
Additive and {{{}@Internal{}}}. {{parseJson(String)}} stays and is
reimplemented on top of the new method by adapting the existing Jackson parser
to {{{}VariantJsonSource{}}}. {{PARSE_JSON}} and {{TRY_PARSE_JSON}} behavior is
unchanged.
*Verifying this change*
* Existing {{PARSE_JSON}} / {{TRY_PARSE_JSON}} and
{{BinaryVariantInternalBuilder}} tests pass unchanged, proving the String path
still behaves identically.
* Add a test that builds a {{Variant}} from a non-Jackson
{{VariantJsonSource}} implementation and asserts it is byte-for-byte equal to
the Variant produced by {{parseJson(String)}} for the same document, across
objects, arrays, and all scalar and numeric forms.
> Allow building a Variant from a token source, without a String round-trip
> -------------------------------------------------------------------------
>
> Key: FLINK-40218
> URL: https://issues.apache.org/jira/browse/FLINK-40218
> Project: Flink
> Issue Type: Improvement
> Components: API / Type Serialization System
> Reporter: Ramin Gharib
> Assignee: Ramin Gharib
> Priority: Major
> Labels: pull-request-available
>
> The only way to build a \{{Variant}} from JSON today is from a \{{String}}.
> \{{PARSE_JSON}} and the internal
> \{{BinaryVariantInternalBuilder.parseJson(String)}} both create a Jackson
> parser over the text and walk its tokens in \{{buildJson}}.
> A caller that already holds parsed JSON cannot use that work. It must write
> its data back to a \{{String}}, which flink-core then parses a second time:
> {code}
> bytes --(caller parses)--> tree --(toString)--> String --(flink-core parses
> again)--> Variant
> {code}
> Two things keep callers on this path:
> # *There is no entry point that takes a parser.* \{{parseJson(JsonParser,
> boolean)}} exists but is private.
> # *Jackson types do not cross shading boundaries.* The builder uses Flink's
> shaded \{{org.apache.flink.shaded.jackson2...JsonParser}}. A caller with its
> own, unshaded Jackson cannot pass in its parser or tree. A \{{String}} is the
> only type both sides share.
> *Proposed change*
> Two entry points on the \{{@Internal}} \{{BinaryVariantInternalBuilder}}:
> ||The caller holds||Entry point||
> |Flink's shaded Jackson \{{JsonParser}}|\{{parseJson(JsonParser, boolean)}},
> now public|
> |JSON in another form, such as a tree from its own Jackson|walk it with the
> existing \{{append*}}, \{{addKey}} and \{{finishWritingObject}} methods, and
> call the new \{{appendJsonNumber(String)}} for numbers|
> {\{parseJson(JsonParser, boolean)}} expects the parser on the value's first
> token and leaves it on the value's last token, so a format can read a VARIANT
> field in the middle of a record.
> {\{appendJsonNumber(String)}} reads the literal with the same Jackson factory
> and number code as \{{PARSE_JSON}}. A number is therefore stored byte for
> byte like \{{PARSE_JSON}} stores it: the smallest integer, else a decimal,
> else a finite double. Anything that is not exactly one JSON number fails.
> A walker over an unshaded Jackson tree then looks li
> {code:java}
> case NUMBER:
> if (node.isIntegralNumber() && node.canConvertToLong()) {
>
> builder.appendNumeric(node.longValue());
> } else {
>
> builder.appendJsonNumber(node.asText());
> }
> \{code}
>
> *Design notes*
> * *Push instead of a token-source interface.* An earlier proposal added a
> pull-based \{{VariantJsonSource}} interface that builder reads from. It is
> replaced by one push methocers push into the builder too, such
> as\{{ToVariantConverter}} for \{{CAST(... AS VARIANT)}}. One method is a
> smaller surface than an interface with an enum. {{PARSkeeps its own code path
> unchanged.
> * *A String literal, not a \{{Number}}.* \{{PARSE_JSON}} picks the type from
> how a number is written, for example an exponenDOUBLE. A \{{BigDecimal}}
> cannot tell \{{1e-7}} from \{E_JSON}} stores as DOUBLE and DECIMAL. The text
> keeps that information.
> * *No hand-written number
> grammar.* \{{appendJsonNume the literal, so it rejects {{+5}},
> \{{0x1p4}},\{{1.5f}} and \{{1,5}} exactly like \{{PARSE_JSON}} does.
>
> *Scope*
>
> Moving the \{{json}} format onto the new parser entry point is split into
> FLINK-XXXXX. That change is not behavior-neutral: format rounds VARIANT
> floats through a \{{double}} to0}} becomes the string \{{"Infinity"}}. It
> needs itsown release note.
>
> *Compatibility*
>
> Additive and \{{@Internal}}. \{{PARSE_JSON}}, \{{TRY_PARSE_JSON}} and the
> \{{json}} format do not change. \{{addKey}} now looks once instead of twice.
>
> *Verifying this change*
> * \{{appendJsonNumber}} is byte-for-byte equal to \{{parseJson(String)}} for
> integers at every width, values beyond a long, decimals, exponents and
> surrounding whitespace.
> * It rejects non-JSON literals, other JSON values such as \{{"5"}} or
> \{{[1]}}, trailing content such as \{{5 6}}, and numbers outside the double
> range.
> * A parser read in the middle of a document is left on the value's last
> token.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)