Claus Ibsen created CAMEL-25077:
-----------------------------------

             Summary: camel-xml-io - XML DSL parser and dumper: fix bugs found 
in a deep review
                 Key: CAMEL-25077
                 URL: https://issues.apache.org/jira/browse/CAMEL-25077
             Project: Camel
          Issue Type: Bug
          Components: camel-core
            Reporter: Claus Ibsen


A deep review of the XML DSL parser and dumper (camel-xml-io, 
camel-xml-io-util) found the bugs below. Each one was reproduced against 
4.23.0-SNAPSHOT and has a test that fails without the fix 
(MXParserEdgeCasesTest, XmlStreamReaderTest, DumpModelEdgeCasesTest).

# *Two CDATA sections in a row followed by text duplicate the second one, or 
fail the parser.* <constant><![CDATA[x]]><![CDATA[y]]>z</constant> gave xyyz, 
and the idiom to put ]]> in CDATA followed by a new line 
(<![CDATA[a]]]]><![CDATA[>b]]>) gave a]]>b>b. With a comment in between it 
failed with ArrayIndexOutOfBoundsException. The parser joined the first section 
into its buffer when it read the second one, but still marked it as to be 
joined, so the following text joined it again. This is from the original xpp3 
code.
# *A character reference above U+FFFF is silently truncated.* &#x1F600; gave 
U+F600 instead of the emoji, as the code point was kept in a char. A reference 
that is not a valid character now fails.
# *dumpBeansAsXml produces malformed XML.* The <script> of a bean was written 
inside its start tag, and no value was escaped, so a property value such as a 
url with & gave invalid XML (also in camel-xml-jaxb, which has a copy of that 
code). A bean without a type failed with NullPointerException.
# *The XML and YAML dumpers drop the note of the routes and EIPs* (added in 
CAMEL-22576), as their override of the id/description attributes did not write 
it.
# *An encoding declared on another line of the XML declaration is ignored.* 
<?xml version="1.0"\n encoding="ISO-8859-1"?> was read as UTF-8 (so café was 
read wrongly), as the pattern did not match across lines.

*Not changed (for a later look)*
* Whitespace-only expression text is dropped even with trim="false".
* Each dump and re-parse cycle adds another copy of the xpath namespace.
* An error for an invalid attribute points at the end of the start tag, not 
where it starts.
* With CR-only line endings every line number is 1.
* A UTF-8 BOM that disagrees with the declared encoding: the declaration wins.

_Claude Code on behalf of Claus Ibsen_




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to