[ 
https://issues.apache.org/jira/browse/FOP-2918?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18122098#comment-18122098
 ] 

Jason Harrop commented on FOP-2918:
-----------------------------------

The same split happens on a second path. With 
{{font-selection-strategy="character-by-character"}}, {{TextLayoutManager}} 
also ends the word wherever the per-character font selection picks a different 
font. When the font changes at a surrogate pair, the word is cut between the 
high and the low surrogate, and the fragment is rejected the same way, so no 
PDF is produced:

{noformat}
java.lang.IllegalArgumentException: ill-formed UTF-16 sequence, contains
    isolated high surrogate at end of sequence
  at org.apache.fop.fonts.MultiByteFont.mapCharsToGlyphs(MultiByteFont.java:695)
  at 
org.apache.fop.fonts.MultiByteFont.charSequenceToGlyphSequence(MultiByteFont.java:599)
  at 
org.apache.fop.fonts.MultiByteFont.performSubstitution(MultiByteFont.java:580)
  at org.apache.fop.fonts.Font.performSubstitution(Font.java:471)
  at org.apache.fop.fonts.GlyphMapping.processWordMapping(GlyphMapping.java:141)
  at org.apache.fop.fonts.GlyphMapping.doGlyphMapping(GlyphMapping.java:92)
  at 
org.apache.fop.layoutmgr.inline.TextLayoutManager.processWord(TextLayoutManager.java:982)
  at 
org.apache.fop.layoutmgr.inline.TextLayoutManager.getNextKnuthElements(TextLayoutManager.java:841)
{noformat}

To reproduce: a block with {{font-family="Helvetica, Aegean600"}} and 
{{font-selection-strategy="character-by-character"}} containing U+10300, using 
the Aegean600 font already in FOP's test resources. Any code point outside the 
BMP at which the selected font changes does it; emoji and CJK Extension B are 
the common cases.

The fix is this issue's guard applied to both places the word is ended: the 
bidi-level change in the attached patch, and the font change. A pull request 
against current main follows, titled FOP-2918, with both guards, the 
{{CharUtilities.containsSurrogatePairAt}} correction from the attached patch 
({{>}} to {{>=}}), and tests: 
{{PDFEncodingTestCase.testPDFEncodingWithNonBMPFontCharacterByCharacter}} 
(fails before with the exception above) and 
{{CharUtilitiesTestCase.testContainsSurrogatePairAtWithIsolatedHighSurrogateAtEndOfSequence}}.
 The {{TextLayoutManager}} change reached us through the Metanorma fork of FOP 
(Alexander Dyuzhev, their issue #39), with authorship kept.


> [PATCH] Surrogate pairs not handled in U+10800-U+1083F
> ------------------------------------------------------
>
>                 Key: FOP-2918
>                 URL: https://issues.apache.org/jira/browse/FOP-2918
>             Project: FOP
>          Issue Type: Bug
>          Components: renderer/pdf
>    Affects Versions: 2.4
>         Environment: Windows 10
>            Reporter: Jan Driesen
>            Priority: Major
>         Attachments: 2918.patch, NotoSansCypriot-Regular.ttf, fop.xconf, 
> input.fo
>
>
> Fop is not properly handling surrogate pairs for characters in Unicode Block 
> 'Cypriot Syllabary' when rendering PDF.
> It tries to resolve the individual surrogate entities. This results in errors 
> saying the glyphs cannot be found.
> The attached test shows a font that supports characters in this range, and an 
> FO file holding the surrogate characters to be rendered.
> Similar issues arise with fonts "MPH 2b Damas" 
> ([https://fedoraproject.org/wiki/MPH_2B_Damase_fonts]) and "Segoe UI 
> Historic" 
> ([https://docs.microsoft.com/en-us/typography/font-list/segoe_ui_historic),] 
> but the error may differ. [I am unsure whether licensing allows me to add 
> these)
> Some fonts (Damas & Noto) result in a "String index out of range". Other 
> fonts (Segoe) deliver a "ill-formed UTF-16 sequence, contains isolated high 
> surrogate at end of sequence" FOPException.
> We expected this to work thanks to FOP-1969 (fop 2.3).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to