allthingssecurity opened a new pull request, #27035: URL: https://github.com/apache/camel/pull/27035
# Description [CAMEL-25126](https://issues.apache.org/jira/browse/CAMEL-25126) The syslog data format and the `SyslogMessage` type converter corrupt any text outside US-ASCII. `SyslogConverter.parseMessage` turns each byte into one character with `(char) (byteBuffer.get() & 0xff)`, which is an ISO-8859-1 decoding, and `SyslogDataFormat.unmarshal` feeds it `body.getBytes()` of a String that it has just decoded with the exchange charset. On a JVM whose default charset is UTF-8 (every JVM since Java 18, unless configured otherwise) a UTF-8 message comes out as mojibake: ``` from("netty:udp://0.0.0.0:10514?sync=false&allowDefaultCodec=false").unmarshal().syslog() datagram (UTF-8): <14>Oct 1 10:00:00 host café Grüße log message: café GrüÃe ``` (`ß` is the bytes `C3 9F`, so after the `Ã` there is the invisible control character U+009F.) RFC 5424 specifies UTF-8 for MSG (section 6.4, MSG-UTF8 = BOM UTF-8-STRING) and for structured data values (section 6.3.3). The BOM ends up as `` at the start of the log message. The `encoding` option of the netty or mina endpoint does not help, because the data format re-encodes the text with the JVM default charset before it parses it, and `unmarshal(marshal(m))` does not give `m` back for any non-ASCII text. This change: - `SyslogConverter.parseMessage(byte[], Charset)` (new) decodes the bytes of each field with the charset. The loops that read the fields are unchanged: they still read one char per byte, and each field is turned back into its bytes and decoded when it is set. `parseMessage(byte[])` uses UTF-8, and so does `toSyslogMessage(String)` for its internal bytes, so a String is parsed without loss. - A MSG that starts with the UTF-8 BOM is decoded as UTF-8, whatever the charset, and the BOM is not part of the log message. - `SyslogDataFormat.unmarshal` parses the bytes as received and decodes them with `ExchangeHelper.getCharset(exchange)`: the `CamelCharsetName` header or property (which the `encoding` option of camel-netty and camel-mina sets), else Camel's default charset, UTF-8. `marshal` writes with the same charset. Compatibility: US-ASCII messages are parsed as before, and the existing tests pass unchanged. Where the current code produced the right text (a JVM whose default charset is ISO-8859-1, with characters in that range) the result is the same. Everything else was wrong before. Routes that repaired the text themselves will see a change, so there is a note in the 4.23 upgrade guide. The loops are only touched at the lines that set the fields. The other syslog parsing problems of the companion issue (no MSG, escaped `]`, NILVALUE timestamp) are fixed in a separate PR that changes the loop conditions, and the two merge without conflict in either order. Tests: - New `SyslogCharsetTest`: the converter for RFC 3164 and RFC 5424 (MSG and a structured data value), `parseMessage` of UTF-8 bytes, a MSG with a BOM (also with ISO-8859-1 as the charset), `parseMessage` with ISO-8859-1, `unmarshal` of UTF-8 bytes and of ISO-8859-1 bytes with `CamelCharsetName`, the marshal/unmarshal round trip of `SyslogMessage` and `Rfc5424SyslogMessage`, and a netty UDP route with `unmarshal().syslog()` that receives a UTF-8 datagram. - Without the change (only a `parseMessage(byte[], Charset)` that ignores the charset, so that the test compiles), 9 of the 10 new tests fail with the mojibake above (for example `café GrüÃe æ¥æ¬` instead of `café Grüße 日本`, `Jürgen` instead of `Jürgen`, and a leading `` with a BOM). The one that passes is `parseMessage` of ISO-8859-1 bytes with ISO-8859-1, because the old per-byte decoding is ISO-8859-1. - With the change, the whole camel-syslog module: 23 tests (13 existing, 10 new), 0 failures, 0 errors, 0 skipped. Found with a Lean 4 model of the byte decoding (the decoded text is longer than the sent text whenever it has a character outside US-ASCII, and a UTF-8 decoding per field round trips every Unicode text), then reproduced with the real classes and end to end with a netty UDP route. # Target - [x] I checked that the commit is targeting the correct branch (Camel 4 uses the `main` branch) # Tracking - [x] If this is a large change, bug fix, or code improvement, I checked there is a [JIRA issue](https://issues.apache.org/jira/browse/CAMEL) filed for the change (usually before you start working on it). # Apache Camel coding standards and style - [x] I checked that each commit in the pull request has a meaningful subject line and body. - [ ] I have run `mvn clean install -DskipTests` locally from root folder and I have committed all auto-generated changes. (I built and tested the affected module, including the formatter and import-sort plugins. I did not run the full root build.) # AI-assisted contributions - [x] If this PR includes AI-generated code, commits have proper co-authorship attribution (e.g., `Co-authored-by` trailers) and the PR description identifies the AI tool used. This PR was prepared with Claude Code (Claude Opus 5.5). The commit carries a `Co-Authored-By` trailer. _Claude Code on behalf of allthingssecurity_ 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
