manner opened a new pull request, #29341:
URL: https://github.com/apache/flink/pull/29341

   ## What is the purpose of the change
   
   This PR adds the `PARSE_XML` and `TRY_PARSE_XML` functions from 
[FLIP-612](https://cwiki.apache.org/confluence/spaces/FLINK/pages/451974259/FLIP-612+Native+XML+Functions+for+Flink+SQL).
 They parse an XML string into a `VARIANT`, so that XML data can be queried the 
same way as JSON data from `PARSE_JSON`. `XML_STRING` follows in a separate PR.
   
   This PR depends on #29323 (FLINK-40832), which validates decimals in the 
variant builder. Until it is merged, three decimal tests in 
`XmlToVariantParserTest` fail.
   
   ## Brief change log
   
   - `XmlToVariantParser` reads the document with the JDK StAX parser into a 
tree of elements and writes it as a variant. Attributes are stored as `@name` 
fields, text as `$`, repeated child elements as arrays, and `#` records the 
document order. `xsi:type` stores the text as a typed value, and 
`xsi:nil="true"` stores the element as `NULL`.
   - `PARSE_XML(xml[, force_array])` and `TRY_PARSE_XML(xml[, force_array])` as 
built-in functions, in the Table API (`parseXml()`, `tryParseXml()`), and in 
the Python API. With `force_array`, every child element is stored as an array, 
so that the shape doesn't depend on how often an element occurs.
   - Internal entities from the DTD are expanded. External entities and 
external DTDs are refused, and entity expansion and nesting depth are limited, 
to prevent XXE and billion laughs attacks.
   - Docs for both functions, and a description of the XML mapping in the 
`VARIANT` data type docs.
   - A separate commit moves `microsSinceEpoch` and `nanosSinceEpoch` from 
`BinaryVariantBuilder` to `BinaryVariantUtil`, so that the parser can reuse 
them.
   
   A few details that the FLIP doesn't spell out:
   - `xsi:type` ignores the prefix, so `xs:int`, `xsd:int`, and `int` are the 
same type. Documents almost always use a prefix, since the built-in types are 
in the XML Schema namespace.
   - `xsi:nil="true"` takes precedence over the content. On an element with 
other attributes, the attributes are kept, and `@xsi:nil` and `@xsi:type` are 
kept with them.
   - XML 1.1 documents are rejected.
   - `force_array` accepts any `BOOLEAN` expression, like the second argument 
of `PARSE_JSON`. A `NULL` value returns `NULL`.
   
   ## Verifying this change
   
   This change added tests and can be verified as follows:
   
   - `XmlToVariantParserTest` covers the mapping rules, `xsi:type` and 
`xsi:nil`, `force_array`, the normalization of text, comments, CDATA and 
namespaces, and the security limits (external entities and DTDs, number and 
size of entity expansions, nesting depth, XML 1.1).
   - `XmlFunctionsITCase` covers SQL and the Table API for both functions, 
including `NULL` input, invalid documents, `force_array` from a literal, a 
column, and `NULL`, constant folding, and casts of the result.
   - `test_expression.py` covers the Python API.
   
   ## Does this pull request potentially affect one of the following parts:
   
     - Dependencies (does it add or upgrade a dependency): **no**
     - The public API, i.e., is any changed class annotated with 
`@Public(Evolving)`: **yes**, new methods in `BaseExpressions`
     - The serializers: **no**
     - The runtime per-record code paths (performance sensitive): **yes**, the 
new functions parse each record, but existing code paths are unchanged
     - Anything that affects deployment or recovery: JobManager (and its 
components), Checkpointing, Kubernetes/Yarn, ZooKeeper: **no**
     - The S3 file system connector: **no**
   
   ## Documentation
   
     - Does this pull request introduce a new feature? **yes**
     - If yes, how is the feature documented? **docs** and **JavaDocs**
   
   ---
   
   ##### Was generative AI tooling used to co-author this PR?
   
   - [X] Yes (please specify the tool below)
   
   Generated-by: Claude Code (Claude Opus 5.5)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to