davsclaus commented on code in PR #27035:
URL: https://github.com/apache/camel/pull/27035#discussion_r4130731896


##########
components/camel-syslog/src/main/java/org/apache/camel/component/syslog/SyslogConverter.java:
##########
@@ -142,10 +147,25 @@ public static String toString(SyslogMessage message) {
 
     @Converter
     public static SyslogMessage toSyslogMessage(String body) {
-        return parseMessage(body.getBytes());
+        return parseMessage(body.getBytes(StandardCharsets.UTF_8), 
StandardCharsets.UTF_8);
     }
 
+    /**
+     * Parses a syslog message whose text is encoded in UTF-8.
+     */
     public static SyslogMessage parseMessage(byte[] bytes) {
+        return parseMessage(bytes, StandardCharsets.UTF_8);
+    }
+
+    /**
+     * Parses a syslog message.
+     *
+     * @param  bytes   the message
+     * @param  charset the charset of the text fields. A MSG that starts with 
the UTF-8 byte order mark is always
+     *                 decoded as UTF-8, without the byte order mark (MSG-UTF8 
of RFC 5424).
+     * @return         the parsed message
+     */
+    public static SyslogMessage parseMessage(byte[] bytes, Charset charset) {

Review Comment:
   Optional: fields are split on raw ASCII bytes (`[`, `]`, `"`, `\`, `=`, 
space) and decoded afterwards, so only ASCII-compatible charsets work. UTF-16, 
EBCDIC or Shift_JIS (whose trail bytes can be `[`, `\` or `]`) would split 
wrongly. Worth saying "ASCII-compatible charset" in the javadoc here and on 
`decode`.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to