Claus Ibsen created CAMEL-25077:
-----------------------------------
Summary: camel-xml-io - XML DSL parser and dumper: fix bugs found
in a deep review
Key: CAMEL-25077
URL: https://issues.apache.org/jira/browse/CAMEL-25077
Project: Camel
Issue Type: Bug
Components: camel-core
Reporter: Claus Ibsen
A deep review of the XML DSL parser and dumper (camel-xml-io,
camel-xml-io-util) found the bugs below. Each one was reproduced against
4.23.0-SNAPSHOT and has a test that fails without the fix
(MXParserEdgeCasesTest, XmlStreamReaderTest, DumpModelEdgeCasesTest).
# *Two CDATA sections in a row followed by text duplicate the second one, or
fail the parser.* <constant><![CDATA[x]]><![CDATA[y]]>z</constant> gave xyyz,
and the idiom to put ]]> in CDATA followed by a new line
(<![CDATA[a]]]]><![CDATA[>b]]>) gave a]]>b>b. With a comment in between it
failed with ArrayIndexOutOfBoundsException. The parser joined the first section
into its buffer when it read the second one, but still marked it as to be
joined, so the following text joined it again. This is from the original xpp3
code.
# *A character reference above U+FFFF is silently truncated.* 😀 gave
U+F600 instead of the emoji, as the code point was kept in a char. A reference
that is not a valid character now fails.
# *dumpBeansAsXml produces malformed XML.* The <script> of a bean was written
inside its start tag, and no value was escaped, so a property value such as a
url with & gave invalid XML (also in camel-xml-jaxb, which has a copy of that
code). A bean without a type failed with NullPointerException.
# *The XML and YAML dumpers drop the note of the routes and EIPs* (added in
CAMEL-22576), as their override of the id/description attributes did not write
it.
# *An encoding declared on another line of the XML declaration is ignored.*
<?xml version="1.0"\n encoding="ISO-8859-1"?> was read as UTF-8 (so café was
read wrongly), as the pattern did not match across lines.
*Not changed (for a later look)*
* Whitespace-only expression text is dropped even with trim="false".
* Each dump and re-parse cycle adds another copy of the xpath namespace.
* An error for an invalid attribute points at the end of the start tag, not
where it starts.
* With CR-only line endings every line number is 1.
* A UTF-8 BOM that disagrees with the declared encoding: the declaration wins.
_Claude Code on behalf of Claus Ibsen_
--
This message was sent by Atlassian Jira
(v8.20.10#820010)