allthingssecurity opened a new pull request, #27035:
URL: https://github.com/apache/camel/pull/27035

   # Description
   
   [CAMEL-25126](https://issues.apache.org/jira/browse/CAMEL-25126)
   
   The syslog data format and the `SyslogMessage` type converter corrupt any 
text outside US-ASCII. `SyslogConverter.parseMessage` turns each byte into one 
character with `(char) (byteBuffer.get() & 0xff)`, which is an ISO-8859-1 
decoding, and `SyslogDataFormat.unmarshal` feeds it `body.getBytes()` of a 
String that it has just decoded with the exchange charset. On a JVM whose 
default charset is UTF-8 (every JVM since Java 18, unless configured otherwise) 
a UTF-8 message comes out as mojibake:
   
   ```
   
from("netty:udp://0.0.0.0:10514?sync=false&allowDefaultCodec=false").unmarshal().syslog()
   
   datagram (UTF-8):  <14>Oct  1 10:00:00 host café Grüße
   log message:       café Grüße
   ```
   
   (`ß` is the bytes `C3 9F`, so after the `Ã` there is the invisible control 
character U+009F.)
   
   RFC 5424 specifies UTF-8 for MSG (section 6.4, MSG-UTF8 = BOM UTF-8-STRING) 
and for structured data values (section 6.3.3). The BOM ends up as `` at the 
start of the log message. The `encoding` option of the netty or mina endpoint 
does not help, because the data format re-encodes the text with the JVM default 
charset before it parses it, and `unmarshal(marshal(m))` does not give `m` back 
for any non-ASCII text.
   
   This change:
   - `SyslogConverter.parseMessage(byte[], Charset)` (new) decodes the bytes of 
each field with the charset. The loops that read the fields are unchanged: they 
still read one char per byte, and each field is turned back into its bytes and 
decoded when it is set. `parseMessage(byte[])` uses UTF-8, and so does 
`toSyslogMessage(String)` for its internal bytes, so a String is parsed without 
loss.
   - A MSG that starts with the UTF-8 BOM is decoded as UTF-8, whatever the 
charset, and the BOM is not part of the log message.
   - `SyslogDataFormat.unmarshal` parses the bytes as received and decodes them 
with `ExchangeHelper.getCharset(exchange)`: the `CamelCharsetName` header or 
property (which the `encoding` option of camel-netty and camel-mina sets), else 
Camel's default charset, UTF-8. `marshal` writes with the same charset.
   
   Compatibility: US-ASCII messages are parsed as before, and the existing 
tests pass unchanged. Where the current code produced the right text (a JVM 
whose default charset is ISO-8859-1, with characters in that range) the result 
is the same. Everything else was wrong before. Routes that repaired the text 
themselves will see a change, so there is a note in the 4.23 upgrade guide.
   
   The loops are only touched at the lines that set the fields. The other 
syslog parsing problems of the companion issue (no MSG, escaped `]`, NILVALUE 
timestamp) are fixed in a separate PR that changes the loop conditions, and the 
two merge without conflict in either order.
   
   Tests:
   - New `SyslogCharsetTest`: the converter for RFC 3164 and RFC 5424 (MSG and 
a structured data value), `parseMessage` of UTF-8 bytes, a MSG with a BOM (also 
with ISO-8859-1 as the charset), `parseMessage` with ISO-8859-1, `unmarshal` of 
UTF-8 bytes and of ISO-8859-1 bytes with `CamelCharsetName`, the 
marshal/unmarshal round trip of `SyslogMessage` and `Rfc5424SyslogMessage`, and 
a netty UDP route with `unmarshal().syslog()` that receives a UTF-8 datagram.
   - Without the change (only a `parseMessage(byte[], Charset)` that ignores 
the charset, so that the test compiles), 9 of the 10 new tests fail with the 
mojibake above (for example `café Grüße 日本` instead of `café Grüße 日本`, 
`Jürgen` instead of `Jürgen`, and a leading `` with a BOM). The one that 
passes is `parseMessage` of ISO-8859-1 bytes with ISO-8859-1, because the old 
per-byte decoding is ISO-8859-1.
   - With the change, the whole camel-syslog module: 23 tests (13 existing, 10 
new), 0 failures, 0 errors, 0 skipped.
   
   Found with a Lean 4 model of the byte decoding (the decoded text is longer 
than the sent text whenever it has a character outside US-ASCII, and a UTF-8 
decoding per field round trips every Unicode text), then reproduced with the 
real classes and end to end with a netty UDP route.
   
   # Target
   
   - [x] I checked that the commit is targeting the correct branch (Camel 4 
uses the `main` branch)
   
   # Tracking
   - [x] If this is a large change, bug fix, or code improvement, I checked 
there is a [JIRA issue](https://issues.apache.org/jira/browse/CAMEL) filed for 
the change (usually before you start working on it).
   
   # Apache Camel coding standards and style
   
   - [x] I checked that each commit in the pull request has a meaningful 
subject line and body.
   - [ ] I have run `mvn clean install -DskipTests` locally from root folder 
and I have committed all auto-generated changes.
     (I built and tested the affected module, including the formatter and 
import-sort plugins. I did not run the full root build.)
   
   # AI-assisted contributions
   
   - [x] If this PR includes AI-generated code, commits have proper 
co-authorship attribution (e.g., `Co-authored-by` trailers) and the PR 
description identifies the AI tool used.
     This PR was prepared with Claude Code (Claude Opus 5.5). The commit 
carries a `Co-Authored-By` trailer.
   
   _Claude Code on behalf of allthingssecurity_
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to