[ 
https://issues.apache.org/jira/browse/FOP-3346?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18123942#comment-18123942
 ] 

Joao Goncalves commented on FOP-3346:
-------------------------------------

Can you add an FO to replicate the issue?

> ToUnicode selectors are off by one after every supplementary-plane character, 
> so the text after an emoji or a mathematical letter extracts wrongly
> --------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: FOP-3346
>                 URL: https://issues.apache.org/jira/browse/FOP-3346
>             Project: FOP
>          Issue Type: Bug
>          Components: renderer/pdf
>    Affects Versions: 2.11
>            Reporter: Jason Harrop
>            Priority: Minor
>
> CIDSubset.getChars builds a char\[] with StringBuilder.appendCodePoint, so a 
> supplementary-plane character occupies two slots. PDFToUnicodeCMap then 
> derives the character selector from the array position. Every selector after 
> the pair is therefore written one too high.
> Measured with the 2.11 command line, DejaVu Math TeX Gyre, the text "A𝐀BZ": 
> the content stream uses selectors 3 4 5 6, and the CMap says
> {noformat}
> <0003> <0041>  <0004> <d835dc00>  <0006> <0042>  <0007> <005a>
> {noformat}
> Selector 5, the B, has no entry, and selector 6, the Z, is published as B. 
> pdftotext extracts "A𝐀 B"; pdf.js and PDFium give the missing selector as 
> U+0005.
> PDFToUnicodeCMapTestCase.surrogatePairTest pins the drift: it expects the 
> entry after the pair at 0x63 to be 0x65.
> h3. Fix
> Build the CMap from one destination per selector (a String, so a surrogate 
> pair is one entry of length two) rather than from a positional char\[]. The 
> range logic then needs no surrogate special cases: an entry may join a 
> bfrange when it is one code point, and two entries are consecutive when their 
> code points are and their selectors share a 256 block. Expectations in 
> surrogatePairTest, surrogatePairRangeTest, surrogatePairsRangeTest and 
> rangeSizeSurrogateTest change accordingly; the last also used low surrogates 
> that ran past U+DFFF and now starts at U+DC00.
> After the fix the same file extracts as "A𝐀BZ" in pdftotext, mupdf, pdf.js 
> and PDFium.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to