shashank created CAMEL-25356:
--------------------------------

             Summary: camel-cm-sms - messages with <, >, ¤ or a form feed are 
sent as unicode: the GSM 03.38 regex contains HTML entities
                 Key: CAMEL-25356
                 URL: https://issues.apache.org/jira/browse/CAMEL-25356
             Project: Camel
          Issue Type: Bug
          Components: camel-cm-sms
            Reporter: shashank


{{CMConstants.GSM_0338_REGEX}} was copied from a web page together with its 
HTML entities:

{code:java}
"^[A-Za-z0-9 \\r\\n@£$Δ_...!\"#$%&amp;'()*+,\\-./:;&lt;=&gt;?¡¿^{}\\\\\\[~\\]|" 
+ "€¥...]*$"
{code}

In a Java character class {{&lt;}} is the four characters {{&}}, {{l}}, {{t}}, 
{{;}}: {{<}} and {{>}} are not in the class. The currency sign {{¤}} (0x24 of 
the GSM default alphabet) and the form feed of the extension table are missing 
too. {{CMMessage.setUnicodeAndMultipart}} then sends any message with one of 
these characters as unicode ({{<DCS>8</DCS>}}): 70 / 67 characters per part 
instead of 160 / 153, and the maximum number of parts is computed for unicode 
(a 310-character message needs 5 parts instead of 3, and a message of more than 
536 characters exceeds the default maximum of 8 unicode parts, while it needs 4 
GSM parts). GSM 03.38 (3GPP TS 23.038; the mapping 
https://unicode.org/Public/MAPPINGS/ETSI/GSM0338.TXT, cited in the new test) 
has these characters, and the {{CMMessage}} javadoc promises 160 / 153 
characters per part for a message in GSM 7-bit characters.

h3. Reproduction

New {{CMGsm0338Test}}: every character of the GSM basic set and of the 
extension table checked with {{CMUtils.isGsm0338Encodeable}}: main reports 
{{U+00A4 U+003C U+003E}} and {{U+000C}} as not GSM; a 310-character text with 
{{<}} and {{>}} is sent as unicode; Cyrillic, {{ê}} and a backtick are the 
control (not GSM). Two runs on main.

h3. Proposed fix

Write the characters themselves in the class (and drop the duplicated {{$}}). 
Every message that matched before still matches. Module: 55 tests pass.

Found with a Lean 4 model of the character class against the GSM 03.38 table: 
{{main_missing}} computes that exactly {{¤ < >}} and form feed are missing, 
{{main_sound}} that main accepts nothing outside GSM 03.38, {{fix_complete}} / 
{{fix_sound}} that the class of the fix is exactly GSM 03.38, 
{{fix_extends_main}} that no GSM message of main becomes unicode.

Not in scope: the extension characters count as two septets in GSM 03.38; the 
part count uses {{String.length()}} (unchanged).

Affected: 4.14.x, 4.18.x and main (the regex dates from the component, 2016).

Duplicate check (2026-10-04): JIRA component camel-cm-sms (2 issues), "cm-sms" 
with unicode: none. GitHub pull requests "cm-sms": dependency bumps only.

_Filed with Claude Code on behalf of allthingssecurity._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to