[
https://issues.apache.org/jira/browse/CAMEL-25152?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Claus Ibsen resolved CAMEL-25152.
---------------------------------
Fix Version/s: 4.23.0
Resolution: Fixed
The fix is merged on main, so it is in Camel 4.23.0:
* 863c1c173949 CAMEL-25152: camel-bindy - fixed-length marshal pads and clips
fields by code points, like unmarshal
Resolving, as the ticket was not updated when the PR was merged.
_Claude Code on behalf of Claus Ibsen_
> camel-bindy - fixed-length marshal pads and clips fields in UTF-16 chars
> while unmarshal counts code points (or graphemes), so a record with an emoji
> or other non-BMP character is written one pad short and read back shifted
> -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25152
> URL: https://issues.apache.org/jira/browse/CAMEL-25152
> Project: Camel
> Issue Type: Bug
> Components: camel-bindy
> Reporter: shashank
> Priority: Minor
> Fix For: 4.23.0
>
>
> Since CAMEL-14521 (3.1.0) the fixed-length unmarshal cuts a record with
> {{UnicodeHelper}}: a field of {{length = 5}} is 5 code points, or 5 graphemes
> with {{@FixedLengthRecord(countGrapheme = true)}}. Marshal was not changed:
> {{BindyFixedLengthFactory.generateFixedLengthPositionMap}} pads a field with
> {{length - result.length()}} padding characters and clips it with
> {{result.substring(0, length)}}, both in UTF-16 chars. The record length
> check of {{BindyFixedLengthDataFormat.createModel}} also uses
> {{String.length()}}.
> For text in the Basic Multilingual Plane with {{countGrapheme = false}} the
> two counts are the same. A character outside the BMP (emoji, CJK Extension B,
> many historic scripts) is one code point but two UTF-16 chars, and with
> {{countGrapheme = true}} a base letter with a combining mark ({{e}} + U+0301)
> is one grapheme but two chars. Such a field is written one padding character
> short per extra char, and unmarshal of the record then takes the missing
> characters from the next field, so every following field is shifted:
> {noformat}
> model: @FixedLengthRecord(length = 12), fields of length 5, 4, 3 (align R,
> trim = true), values "ok😀", "x", "y"
> marshal today: " ok😀 x y" (4 code points for the first field, 11 in all)
> unmarshal: "ok😀 ", "x ", "y"
> fixed marshal: " ok😀 x y" -> "ok😀", "x", "y"
> the correctly padded record " ok😀 x y" (12 code points) written by
> another system:
> unmarshal today: IllegalArgumentException: Size of the record: 13 is not
> equal to the value provided in the model: 12
> {noformat}
> With {{clip = true}} the clip can also cut a surrogate pair in half and write
> an invalid character. With {{countGrapheme = true}} (added for exactly this
> kind of text) a decomposed {{café}} is read back as {{café }} and the next
> fields are shifted the same way.
> h3. Reproduction
> A unit test ({{BindyFixedLengthMarshalUnicodeTest}}) marshals and unmarshals
> records with {{ok😀}}, two CJK Extension B characters and a decomposed
> {{café}} (with {{countGrapheme = true}}), and clips {{abcd😀f}} to 5; all four
> cases fail on main (three runs) and pass with the fix. A small formal model
> (Lean 4) shows that padding by code points gives every record back whose
> values fit their lengths, that any value with a non-BMP character followed by
> another field is read back with extra characters today, and that for BMP text
> the fix writes exactly what is written today.
> h3. Affected versions
> Unmarshal has counted code points since 3.1.0 (CAMEL-14521) and marshal still
> counts UTF-16 chars in 3.1.0, 4.0.0, 4.22.0 and main. Before 3.1.0 both
> counted UTF-16 chars.
> h3. Proposed fix
> * In {{generateFixedLengthPositionMap}} measure and clip the formatted value
> with {{UnicodeHelper}} and the same method as unmarshal ({{CODEPOINTS}}, or
> {{GRAPHEME}} with {{countGrapheme}}): pad with {{length -
> unicodeResult.length()}} and clip with {{unicodeResult.substring(0, length)}}.
> * In {{createModel}} compare and cut the record length
> ({{ignoreTrailingChars}}) with the same count.
> Text in the BMP without combining marks is written and read exactly as today.
> Duplicate check (2026-09-30): JIRA "bindy" with "grapheme", "unicode",
> "surrogate", "emoji", "codepoint": CAMEL-14521 and CAMEL-14825 (unmarshal
> only, fixed), CAMEL-14085 (bytes vs chars, Won't Fix); nothing about marshal.
> No open pull request touches the fixed-length marshal.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)