[
https://issues.apache.org/jira/browse/CAMEL-25303?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Claus Ibsen resolved CAMEL-25303.
---------------------------------
Resolution: Fixed
> camel-barcode - text outside ISO-8859-1 comes back as '?' (the documented
> default encoding UTF-8 is not used), and the hints added before the route
> starts are dropped
> ----------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25303
> URL: https://issues.apache.org/jira/browse/CAMEL-25303
> Project: Camel
> Issue Type: Bug
> Components: camel-barcode
> Reporter: shashank
> Assignee: shashank
> Priority: Major
> Fix For: 4.23.0
>
>
> h3. 1. Text that ISO-8859-1 cannot represent is lost
> {{BarcodeDataFormat}} gives ZXing no {{EncodeHintType.CHARACTER_SET}}. ZXing
> then writes the text in ISO-8859-1 (QR code, Aztec; PDF417 refuses such
> text), replacing every other character by {{?}}. With the default data format:
> * {{日本語のテキスト}} is read back as {{????????}};
> * {{Привет, мир}} as {{??????, ???}} (QR code and Aztec);
> * PDF417 fails with {{WriterException: Non-encodable character detected: П}}.
> The documentation of the data format lists "encoding: UTF-8" among the
> default values. Text that ISO-8859-1 can represent ({{Grüße aus Köln}}) comes
> back.
> h3. 2. The hints added before the route starts are dropped (regression in
> 4.15.0)
> Since CAMEL-22354 the hints are computed in {{doStart()}} by
> {{optimizeHints()}}, which starts with {{writerHintMap.clear();
> readerHintMap.clear();}}. The documented way to configure ZXing is
> {{addToHintMap(...)}} on the data format instance, which a route does before
> the route (and so the data format) starts; those hints are cleared and never
> used. Before 4.15.0 (CAMEL-22354's fix version; so 4.18.x LTS is affected,
> 4.14.x is not) the hints were computed in the constructors and setters, so
> hints added afterwards were kept (until a later setter call, CAMEL-7870).
> This also means the CHARACTER_SET hint cannot be used as a workaround for 1
> in a route.
> h3. Reproduction
> New {{BarcodeDataFormatCharsetTest}} (marshal and unmarshal through the data
> format API): Japanese and Cyrillic text in a QR code, Cyrillic in an Aztec
> code and in a PDF417 code: fail on main; {{testHintsAddedBeforeStart}}: a
> {{CHARACTER_SET}} and a {{PURE_BARCODE}} hint added before {{start()}} are
> gone after start ({{expected: <UTF-8> but was: <null>}});
> {{testCharacterSetHintIsUsed}}: with a {{CHARACTER_SET=UTF-8}} hint added
> before start, {{Grüße}} is still written in ISO-8859-1 (byte segment of 5
> bytes instead of 7). Control: ASCII and ISO-8859-1 text round trip on main
> and with the fix.
> h3. Proposed fix
> * When no {{CHARACTER_SET}} hint is given and the text has a character that
> ISO-8859-1 cannot represent, encode it as UTF-8 (ZXing writes an ECI, which
> the readers use). Text that ISO-8859-1 can represent is encoded exactly as
> today (no hint, so no ECI: ZXing 3.5.4's QR encoder appends an ECI only when
> a {{CHARACTER_SET}} hint is present), so existing barcodes do not change.
> This applies to the formats whose ZXing writers take {{CHARACTER_SET}}: QR
> code, Aztec, PDF417; the Data Matrix writer uses it only with
> {{DATA_MATRIX_COMPACT}}, the linear formats never (no change for those).
> * Keep the hints added with {{addToHintMap}} apart and apply them on top of
> the default hints whenever the hints are computed ({{removeFromHintMap}}
> removes them from both), so CAMEL-7870's re-optimization still drops the
> defaults of another format but no longer the user's hints.
> * The documentation's default encoding is corrected (ISO-8859-1, or UTF-8
> when needed; a CHARACTER_SET hint overrides it; the formats that use it);
> catalog copy updated.
> With the fix the 7 new tests and the module pass (33 tests).
> Found with a Lean 4 model of the text encoding and of the hint maps: the
> property "unmarshal(marshal(text)) = text" fails for every text with a
> character outside ISO-8859-1 (theorem {{main_loses}}) and holds with the fix
> for every Unicode text ({{fix_roundtrip}}, using the UTF-8 round trip proved
> for CAMEL-25126), with exactly today's bytes for ISO-8859-1 text
> ({{fix_same_latin1}}); "a hint added before start is in force after start"
> fails today ({{main_drops_hint}}) and holds with the fix ({{fix_keeps_hint}}).
> Affected: 1 in every version (4.14.x, 4.18.x, main; the code is the same
> since the component was added); 2 since 4.15.0 (4.18.x and main; checked on
> camel-4.14.x and the camel-4.15.0 tag).
> Duplicate check (2026-10-03): JIRA component camel-barcode (CAMEL-15535,
> CAMEL-7871, CAMEL-7870, CAMEL-7866) and text "barcode" with "UTF-8" /
> "charset" / "encoding" / "hint": none about the character set or the dropped
> hints. GitHub pull requests "barcode": none on this (open #22358 only
> replaces {{getOut()}} in {{readImage}}).
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)