Jason Harrop created FOP-3346:
---------------------------------

             Summary: ToUnicode selectors are off by one after every 
supplementary-plane character, so the text after an emoji or a mathematical 
letter extracts wrongly
                 Key: FOP-3346
                 URL: https://issues.apache.org/jira/browse/FOP-3346
             Project: FOP
          Issue Type: Bug
          Components: renderer/pdf
    Affects Versions: 2.11
            Reporter: Jason Harrop


CIDSubset.getChars builds a char\[] with StringBuilder.appendCodePoint, so a 
supplementary-plane character occupies two slots. PDFToUnicodeCMap then derives 
the character selector from the array position. Every selector after the pair 
is therefore written one too high.

Measured with the 2.11 command line, DejaVu Math TeX Gyre, the text "A𝐀BZ": the 
content stream uses selectors 3 4 5 6, and the CMap says

{noformat}
<0003> <0041>  <0004> <d835dc00>  <0006> <0042>  <0007> <005a>
{noformat}

Selector 5, the B, has no entry, and selector 6, the Z, is published as B. 
pdftotext extracts "A𝐀 B"; pdf.js and PDFium give the missing selector as 
U+0005.

PDFToUnicodeCMapTestCase.surrogatePairTest pins the drift: it expects the entry 
after the pair at 0x63 to be 0x65.

h3. Fix

Build the CMap from one destination per selector (a String, so a surrogate pair 
is one entry of length two) rather than from a positional char\[]. The range 
logic then needs no surrogate special cases: an entry may join a bfrange when 
it is one code point, and two entries are consecutive when their code points 
are and their selectors share a 256 block. Expectations in surrogatePairTest, 
surrogatePairRangeTest, surrogatePairsRangeTest and rangeSizeSurrogateTest 
change accordingly; the last also used low surrogates that ran past U+DFFF and 
now starts at U+DC00.

After the fix the same file extracts as "A𝐀BZ" in pdftotext, mupdf, pdf.js and 
PDFium.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to