Hi all,

I'm wondering if defaults shouldn't be the other way around? First, we should aim for an `a.b.c` access syntax as much as possible. An array syntax `a[1].b[1].c` is very cumbersome. We can also add a XML_PATH function in the future that uses a string path e.g. XML_PATH("a.b.c") and works dynamically with either fields or arrays under the same parent. JSON_VALUE has a similar logic if I remember correctly.

Alternatively, I could enable `a[1].b[1].c` with an additional parameter once we collected enough feedback.

Cheers,
Timo


On 18.09.26 10:19, Fabian Hüske via dev wrote:
Hi Moritz, hi all,

The idea would be to make ARRAY the default for all elements, so the
structure stays the same no matter how many elements are actually there.
The downside is that for simple XMLs, where most elements only ever appear
once, you'd need to add [1] everywhere just to access them.

One option would be to start with ARRAY by default and see how it goes
based on feedback. If it turns out to be too cumbersome for simple XMLs,
we
could add an opt-out parameter later, that would be a backward-compatible
addition so it doesn't need to be solved now.
Does this make sense, or is there a better way to approach this?

Sounds good to me, thanks!

Best, Fabian

On Thu, Sep 17, 2026 at 5:30 PM Krishna Sai <[email protected]> wrote:

Hi Moritz,

Nice FLIP. The # field is a neat solution, and I checked on master that
Variant does sort object keys, so it is doing necessary work.

One small question: what does PARSE_XML do with an xmlns declaration in
the input? The parser is not namespace-aware, so xmlns:xsi is an
attribute, and by the @name rule it would land as "@xmlns:xsi". If
XML_STRING writes that back and also adds its own xmlns:xsi declaration
for xsi:nil, the root would carry that attribute twice, which is not
well-formed, so that Variant would not round trip.

   <root xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance";>
     <a xsi:nil="true"/>
   </root>

Dropping xmlns at parse time would instead leave other prefixes
undeclared. Maybe I am misreading the mapping, but a line either way
might help.

Best,
Sai Krishna Sepuri

On Thu, Sep 17, 2026 at 7:51 PM Moritz Manner <[email protected]>
wrote:

Hi Fabian,

Good point, and I agree this is a real problem. A consistent VARIANT
structure across documents with the same schema is important, especially
for streaming SQL where you don't want CASE WHEN everywhere, or a job
failing because it tries to access an array that doesn't exist.
Initially,
the structure optimized for queryability, i.e. a single element stays a
single node instead of an array. But I think consistency should probably
win here.

The idea would be to make ARRAY the default for all elements, so the
structure stays the same no matter how many elements are actually there.
The downside is that for simple XMLs, where most elements only ever
appear
once, you'd need to add [1] everywhere just to access them.

One option would be to start with ARRAY by default and see how it goes
based on feedback. If it turns out to be too cumbersome for simple XMLs,
we
could add an opt-out parameter later, that would be a backward-compatible
addition so it doesn't need to be solved now.
Does this make sense, or is there a better way to approach this?

Best,
Moritz

On Thu, 17 Sept 2026 at 15:07, Fabian Hüske via dev <
[email protected]>
wrote:

Hi everyone and thanks for the proposal Moritz,

The FLIP reads very well and I agree (mostly) with the design.

I have only one concern, which is about the representation of repeated
elements as arrays.
I think it is very important that a batch of XML documents with the
same
structure converts into an identical VARIANT structure.
If I understand the current proposal correctly, this would not be the
case
for two documents: one with a single element and one with a repeated
element.
The first document would have a single node, while the second document
would have an array.
If you want to access the nodes, you would need a cumbersome CASE WHEN
construct to handle both cases.
It would be better to convert both documents into a VARIANT containing
an
ARRAY (one only holding a single element).

Since deriving the right conversion strategy from a single XML document
isn't possible, could we add an optional argument to pass conversion
hints?
There might be other solutions to this problem.

IMO, it would also be fine to continue with this FLIP and address this
issue later, but having a proposal for it would be good.
Do you have any thoughts on this, Moritz?

Best, Fabian


On Mon, Sep 14, 2026 at 11:37 AM Timo Walther <[email protected]>
wrote:

Hi Moritz,

this is an excellent design document with the right level of detail.
I
really like the design, it is a nice mixture of concerns. The result
of
PARSE_XML looks and feels like a VARIANT while the XML content is
preserved in a lossless fashion and all attributes and text is
accessible. It is a nice combination of what other vendors offer
(modulo
the historical legacy they have to deal with).

Looking at various UDF implementations, XML handling is by far the
most
important reason for the need of a custom function. So +1 for this
FLIP.

Thanks,
Timo

On 11.09.26 14:54, Moritz Manner wrote:
Hey everyone,

I'd like to start a discussion on FLIP-612: Native XML Functions
for
Flink
SQL [1].

The goal is native XML parsing in Flink SQL, parsing XML into a
VARIANT
instead of requiring custom UDFs.

The FLIP proposes three functions. PARSE_XML and TRY_PARSE_XML
turn an
XML
string into a VARIANT, and XML_STRING turns a VARIANT back into an
XML
string.
This mirrors the existing JSON functions PARSE_JSON,
TRY_PARSE_JSON,
and
JSON_STRING.

Once XML is a VARIANT, you query it the same way as the output of
PARSE_JSON:
with the variant accessors (variant.key, variant['key'],
variant[index])
and a CAST to the type you need.

The main part is the XML→VARIANT mapping: attributes and text go
into
@/$
fields, and since Variant objects store fields sorted by key rather
than
in
document order, a # field is used to recover the original order
between
differently-named siblings. Details and examples are in the FLIP.

It's a small, additive API, no new SQL grammar or Table API methods
needed.

Looking forward to your thoughts!

Best,
Moritz

[1]



https://urldefense.com/v3/__https://cwiki.apache.org/confluence/spaces/FLINK/pages/451974259/FLIP-612*Native*XML*Functions*for*Flink*SQL__;KysrKysr!!Ayb5sqE7!pusFK_SVIXLRWBYFy7nKxC-f5N7Uj1ePbB2T6nFANBnqs-0S_zubmv9UygMYHTQQ1LFe9LUt4tGSVL3srSc$






Reply via email to