[
https://issues.apache.org/jira/browse/CAMEL-25217?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Claus Ibsen resolved CAMEL-25217.
---------------------------------
Fix Version/s: 4.23.0
Resolution: Fixed
The fix is merged on main, so it is in Camel 4.23.0:
* 8b451798f592 CAMEL-25217: camel-platform-http-vertx - write a String response
in the charset of the Content-Type
* bc4afdf796b1 CAMEL-25217: fix the Javadoc of responseCharset
Resolving, as the ticket was not updated when the PR was merged.
_Claude Code on behalf of Claus Ibsen_
> camel-platform-http-vertx - a String response body is always written as
> UTF-8, even when the response Content-Type declares another charset (for
> example charset=ISO-8859-1), so the client decodes it wrongly
> --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25217
> URL: https://issues.apache.org/jira/browse/CAMEL-25217
> Project: Camel
> Issue Type: Bug
> Components: camel-platform-http-vertx
> Reporter: shashank
> Priority: Minor
> Fix For: 4.23.0
>
>
> {{VertxPlatformHttpSupport.toHttpResponse}} copies the {{Content-Type}} of
> the message, with its {{charset}} parameter, to the response.
> {{writeResponse}} then writes a {{String}} body with
> {code:java}
> } else if (body instanceof String string) {
> ctx.end(string);
> {code}
> and Vert.x encodes a {{String}} as UTF-8, whatever the declared charset. A
> client that decodes the response with the charset of the {{Content-Type}} (as
> it should, RFC 9110 8.3.1) gets mojibake for every non-ASCII character. Two
> common cases:
> * a route that sets {{Content-Type: text/plain; charset=ISO-8859-1}} and a
> {{String}} body ({{setBody(constant(...))}}, {{simple}}, {{transform}}):
> {{Grüße aus Köln}} (14 characters) is sent as 17 bytes of UTF-8 labelled
> ISO-8859-1, and read as {{Grü...}} (every non-ASCII character becomes two);
> * an echo, or any route that keeps the request {{Content-Type}}: a request
> with {{charset=ISO-8859-1}} is read correctly (the consumer sets
> {{CamelCharsetName}} from it), but the reply, labelled with the same
> {{Content-Type}}, is UTF-8.
> {{byte[]}}, {{InputStream}} and {{Buffer}} bodies are not affected, and
> neither is ASCII text, which is why the existing tests pass. The type
> converter {{VertxBufferConverter.toBuffer(String, Exchange)}} already uses
> the charset of the message {{Content-Type}} first, then {{CamelCharsetName}}
> (CAMEL-16282); the {{String}} branch of {{writeResponse}} bypasses it.
> (camel-servlet encodes a {{String}} in the exchange charset,
> {{CamelCharsetName}} or UTF-8, and only in non-chunked mode also sets the
> response character encoding to match; so it gets the echo case right but not
> a route that only sets the {{Content-Type}}.)
> h3. Reproduction
> {code:java}
> from("platform-http:/latin1")
> .setHeader(Exchange.CONTENT_TYPE, constant("text/plain;
> charset=ISO-8859-1"))
> .setBody(constant("Grüße aus Köln"));
> from("platform-http:/echo")
> .convertBodyTo(String.class);
> {code}
> {{GET /latin1}} returns {{Content-Type: text/plain; charset=ISO-8859-1}} with
> the UTF-8 bytes of the text (17 bytes instead of 14). {{POST /echo}} with
> {{Content-Type: text/plain; charset=ISO-8859-1}} and the 14 ISO-8859-1 bytes
> returns 17 bytes. A unit test with both routes fails on main; the control
> with {{Content-Type: text/plain}} (no charset, UTF-8 expected) passes. A
> small formal model (Lean 4) shows that for every text with a character
> outside ASCII what a client reads with the declared ISO-8859-1 charset is
> longer than the text, and that writing in the declared charset returns every
> text the charset can represent.
> h3. Proposed fix
> Write the {{String}} in the charset of the response {{Content-Type}} when it
> declares one (and the JVM supports it), otherwise keep UTF-8 as today:
> {code:java}
> } else if (body instanceof String string) {
> // write the text in the charset declared by the response Content-Type
> (vert.x writes a String as UTF-8)
> String charset = responseCharset(ctx);
> if (charset != null) {
> ctx.end(Buffer.buffer(string, charset));
> } else {
> ctx.end(string);
> }
> promise.complete();
> private static String responseCharset(RoutingContext ctx) {
> String contentType = ctx.response().headers().get("Content-Type");
> if (contentType == null) {
> return null;
> }
> String charset = IOHelper.getCharsetNameFromContentType(contentType);
> if (charset == null || charset.isEmpty()) {
> return null;
> }
> try {
> return Charset.isSupported(charset) ? charset : null;
> } catch (IllegalCharsetNameException e) {
> return null;
> }
> }
> {code}
> Only responses that declare a charset other than UTF-8 change, and for them
> the bytes now match the header. A response without a charset parameter stays
> UTF-8 (using {{CamelCharsetName}} there as well, like camel-servlet, would
> change the bytes of responses whose header does not say so, so it is left
> out). With the fix the new test and the whole camel-platform-http-vertx suite
> pass (150 tests, 1 skipped, camel-core built from the same commit). Because
> the bytes of these responses change, the change gets an upgrade guide note (a
> client that ignored the declared charset and read UTF-8 must now use the
> declared one; characters the declared charset cannot represent are written as
> {{?}}).
> Duplicate check (2026-09-30): JIRA "platform-http" with "charset", "encoding"
> or "UTF-8", and "VertxPlatformHttpSupport": CAMEL-16282, CAMEL-16756,
> CAMEL-22125, CAMEL-23320 (Spring Boot starter, binary data), none about a
> String body. No pull request about it.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)