Jason Harrop created FOP-3346:
---------------------------------
Summary: ToUnicode selectors are off by one after every
supplementary-plane character, so the text after an emoji or a mathematical
letter extracts wrongly
Key: FOP-3346
URL: https://issues.apache.org/jira/browse/FOP-3346
Project: FOP
Issue Type: Bug
Components: renderer/pdf
Affects Versions: 2.11
Reporter: Jason Harrop
CIDSubset.getChars builds a char\[] with StringBuilder.appendCodePoint, so a
supplementary-plane character occupies two slots. PDFToUnicodeCMap then derives
the character selector from the array position. Every selector after the pair
is therefore written one too high.
Measured with the 2.11 command line, DejaVu Math TeX Gyre, the text "A𝐀BZ": the
content stream uses selectors 3 4 5 6, and the CMap says
{noformat}
<0003> <0041> <0004> <d835dc00> <0006> <0042> <0007> <005a>
{noformat}
Selector 5, the B, has no entry, and selector 6, the Z, is published as B.
pdftotext extracts "A𝐀 B"; pdf.js and PDFium give the missing selector as
U+0005.
PDFToUnicodeCMapTestCase.surrogatePairTest pins the drift: it expects the entry
after the pair at 0x63 to be 0x65.
h3. Fix
Build the CMap from one destination per selector (a String, so a surrogate pair
is one entry of length two) rather than from a positional char\[]. The range
logic then needs no surrogate special cases: an entry may join a bfrange when
it is one code point, and two entries are consecutive when their code points
are and their selectors share a 256 block. Expectations in surrogatePairTest,
surrogatePairRangeTest, surrogatePairsRangeTest and rangeSizeSurrogateTest
change accordingly; the last also used low surrogates that ran past U+DFFF and
now starts at U+DC00.
After the fix the same file extracts as "A𝐀BZ" in pdftotext, mupdf, pdf.js and
PDFium.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)