shashank created CAMEL-25214:
--------------------------------
Summary: camel-hl7 - HL7DataFormat decodes and encodes messages
with MSH-18 = 8859/6, 8859/7, 8859/8, 8859/9 or 8859/15 in the wrong charset,
and fails for GB 18030-2000 (wrong entries in HL7Charset)
Key: CAMEL-25214
URL: https://issues.apache.org/jira/browse/CAMEL-25214
Project: Camel
Issue Type: Bug
Components: camel-hl7
Reporter: shashank
{{HL7DataFormat}} reads MSH-18 (character set, HL7 table 0211) to choose the
charset for unmarshal and marshal, through {{HL7Charset}}. Some entries of that
enum are wrong:
{code:java}
ISO_8859_5("8859/5", "ISO-8859-5"),
ISO_8859_6("8859/1", "ISO-8859-6"), // the HL7 name should be 8859/6
ISO_8859_7("8859/1", "ISO-8859-7"), // 8859/7
ISO_8859_8("8859/1", "ISO-8859-8"), // 8859/8
ISO_8859_9("8859/1", "ISO-8859-9"), // 8859/9
...
GB_1830_2000("GB 18030-2000", ""), // no Java charset name
{code}
{{getHL7Charset}} returns the first entry whose HL7 name matches, so these four
entries can never be selected ({{8859/1}} finds {{ISO_8859_1}} first), and
{{8859/15}} (in table 0211 since v2.5) has no entry. The result:
||MSH-18||charset used by main||
|{{8859/6}}, {{8859/7}}, {{8859/8}}, {{8859/9}}, {{8859/15}}|the exchange
charset (UTF-8 unless set): Arabic, Greek, Hebrew, Turkish and Latin-9 text is
decoded as mojibake on unmarshal and written in the wrong charset on marshal|
|{{GB 18030-2000}}|{{""}}: unmarshal fails with
{{UnsupportedEncodingException}}, so does marshal|
For example a message with MSH-18 {{8859/7}} and the name {{Παπαδόπουλος}},
written in ISO-8859-7, is unmarshalled as {{Ξ Ξ±ΟΞ±Ξ΄ΟΟΞΏΟλοΟ}}, and the
header {{CamelCharsetName}} says {{UTF-8}}. A patient name with {{Ş}} or {{ğ}}
in an {{8859/9}} message is corrupted the same way.
camel-mllp ({{MllpProtocolConstants.MSH18_VALUES}}) and HAPI
({{ca.uhn.hl7v2.llp.HL7Charsets}}) map these codes to {{ISO-8859-6}} ...
{{ISO-8859-9}}, {{ISO-8859-15}} and {{GB18030}}.
h3. Reproduction
{code:java}
from("direct:unmarshal").unmarshal().hl7(false).to("mock:unmarshal");
String hl7 =
"MSH|^~\\&|MYSENDER|MYSENDERAPP|MYCLIENT|MYCLIENTAPP|200612211200||QRY^A19|1234|P|2.4||||||8859/7"
+
"\rQRD|200612211200|R|I|GetPatient|||1^RD|0101701234^Παπαδόπουλος|DEM||";
template.sendBody("direct:unmarshal", new
ByteArrayInputStream(hl7.getBytes("ISO-8859-7")));
// QRD-8-2 is "Ξ Ξ±ΟΞ±Ξ΄ΟΟΞΏΟλοΟ", CamelCharsetName is UTF-8
{code}
A unit test unmarshals 8859/6, 8859/7, 8859/8, 8859/9, 8859/15 and GB
18030-2000 messages and marshals 8859/7 and GB 18030-2000 messages: all eight
fail on main, the 8859/5 control passes. The entries are the same in 3.0.0,
4.0.0, 4.14.0, 4.18.0 and 4.22.0. A small formal model (Lean 4) of the lookup
checks that main gets exactly these six table 0211 codes wrong, that the four
entries are unreachable, and that the fixed table maps every code on which HAPI
and camel-mllp agree to their charset without changing any code main already
maps correctly.
h3. Proposed fix
{code:java}
ISO_8859_6("8859/6", "ISO-8859-6"),
ISO_8859_7("8859/7", "ISO-8859-7"),
ISO_8859_8("8859/8", "ISO-8859-8"),
ISO_8859_9("8859/9", "ISO-8859-9"),
ISO_8859_15("8859/15", "ISO-8859-15"),
...
GB_1830_2000("GB 18030-2000", "GB18030"),
{code}
The Japanese entries ({{ISO IR14}}, {{ISO IR87}}, {{ISO IR159}}) and {{CNS
11643-1992}} are left as they are: HAPI maps them differently, and Japanese HL7
v2 practice (ISO 2022 code extension, MSH-20) needs more than a table entry.
With the fix the new test and the whole camel-hl7 suite pass (79 tests).
Because messages with these MSH-18 values are now decoded differently (and
{{CamelCharsetName}} changes), the change gets a short upgrade guide note.
Duplicate check (2026-09-30): JIRA "HL7Charset", "MSH-18", "GB 18030",
"8859/7", "8859/9": CAMEL-8079, CAMEL-9867 and CAMEL-22712 (camel-mllp only).
No pull request touches {{HL7Charset}} since it was added.
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)