[
https://issues.apache.org/jira/browse/CAMEL-25043?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Work on CAMEL-25043 started by Claus Ibsen.
-------------------------------------------
> camel-core - Tokenize language: fix bugs found in a deep review
> ---------------------------------------------------------------
>
> Key: CAMEL-25043
> URL: https://issues.apache.org/jira/browse/CAMEL-25043
> Project: Camel
> Issue Type: Bug
> Components: camel-core
> Reporter: Claus Ibsen
> Assignee: Claus Ibsen
> Priority: Major
>
> A deep review of the tokenize language (TokenizeLanguage, and the pair and
> xml iterators in camel-support) found the bugs below. Each one was reproduced
> against 4.23.0-SNAPSHOT and has a test that fails without the fix.
> # *A message without a body fails with a NullPointerException.*
> {{split().tokenize(",")}} of a message without a body (or with a {{source}}
> that has no value) failed with {{NullPointerException: source}} from the
> scanner, where an empty body has no tokens. This is a regression from
> CAMEL-12688 (2018), which removed the fallback to an empty scanner.
> # *skipFirst on an empty body fails with NoSuchElementException.* Without
> {{group}}, skipFirst called {{next()}} on the iterator without checking
> {{hasNext()}}, so an empty file failed the exchange. With {{group}} this was
> already handled (CAMEL-12607).
> # *Pair mode ignores source.*
> {{token("[").endToken("]").source("header:data")}} always read the message
> body. TokenizeLanguage called the overload of {{tokenizePairExpression}}
> without the source, and the overload with a source dropped it too.
> # *skipFirst is ignored in xml mode and in pair mode.* The skip-first wrapper
> was only applied to the plain tokenizer, and xml mode with {{group}} always
> passed {{false}} to the group iterator.
> # *Pair mode with group puts the start token between the pairs.*
> {{token("<a>").endToken("</a>").includeTokens(true).group(2)}} of
> {{<a>1</a><a>2</a><a>3</a>}} gave {{<a>1</a><a><a>2</a>}}. The pairs are now
> joined without a delimiter, as in xml mode, unless {{groupDelimiter}} is set.
> # *A pair with the same start and end token has no tokens.* The end token is
> the scanner delimiter, so a start token that is the same is never found, and
> {{tokenizePair("'", "'")}} silently gave nothing. It is now refused with a
> clear message when the route is created.
> # *xml mode with inheritNamespaceTagName="*" builds broken XML when a comment
> or DOCTYPE comes before the root.* The closing tags were built by treating
> {{<!-- ... -->}} and {{<!DOCTYPE ...>}} as start tags, so each part ended
> with {{</!-->}}.
> # *xml mode with inheritNamespaceTagName="*" fails on a multi-byte character
> before the first token.* The head was cut by the character offset of the
> scanner match, but from the recorded bytes, so an attribute such as
> {{name="æøå"}} failed with StringIndexOutOfBoundsException (or silently lost
> the end of the head).
> # *ValueBuilder.tokenize(token, null, skipFirst) fails when the route
> starts.* It always wrapped the group iterator, also when there is no group,
> which failed with "expression must be specified". The same code in
> MockValueBuilder had the same bug.
> Also: the group delimiter is now encoded in the exchange charset (it was
> encoded in the platform charset and decoded in the exchange charset), and the
> recording input stream of the xml wrap mode no longer drops a NUL byte.
> *Not changed*
> * A token with {{regex=false}} is still compiled as a regular expression, and
> the end token of a pair is not quoted. That is CAMEL-25015, which asks which
> compatible fix the committers prefer.
> * {{SplitterTest.testEmptyBody}} asserted that a request without a body gives
> no out message, which only held because the exchange failed with the
> NullPointerException above. It now asserts that the exchange does not fail
> and has no parts.
> _Claude Code on behalf of Claus Ibsen_
--
This message was sent by Atlassian Jira
(v8.20.10#820010)