[ 
https://issues.apache.org/jira/browse/FOP-3346?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124627#comment-18124627
 ] 

Jason Harrop commented on FOP-3346:
-----------------------------------

Yes. Attached: FOP-3346-selector-drift.fo and fop-test-fonts.xconf, which 
declares two fonts from the test tree (Aegean600.ttf, the one the PDF encoding 
tests use for supplementary-plane characters, and DejaVuLGCSerif.ttf). Copy the 
config to the root of a checkout and run from there:
{noformat}
fop -c fop-test-fonts.xconf -fo FOP-3346-selector-drift.fo -pdf out.pdf
pdftotext out.pdf -
{noformat}
The FO is one block, {{A𐌀BZ}} in Aegean600; U+10300 (Old Italic A) is 
the supplementary-plane character.
On main at 5be8c69b6 the content stream uses selectors 3 4 5 6 and the 
ToUnicode CMap is
{noformat}
<0003> <0041> <0004> <d800df00> <0006> <0042> <0007> <005a>
{noformat}
so selector 5, the B, has no entry and selector 6, the Z, is published as B. 
pdftotext gives {{A𐌀 B}}: the Z is gone.
With the pull request's change the CMap is
{noformat}
<0003> <0041> <0004> <d800df00> <0005> <0042> <0006> <005a>
{noformat}
and pdftotext gives {{A𐌀BZ}}.
The unit test pins the same thing: PDFToUnicodeCMapTestCase.surrogatePairTest 
expected the entry after the pair at 0x65, which was the drift; the PR corrects 
it to 0x64.

> ToUnicode selectors are off by one after every supplementary-plane character, 
> so the text after an emoji or a mathematical letter extracts wrongly
> --------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: FOP-3346
>                 URL: https://issues.apache.org/jira/browse/FOP-3346
>             Project: FOP
>          Issue Type: Bug
>          Components: renderer/pdf
>    Affects Versions: 2.11
>            Reporter: Jason Harrop
>            Priority: Minor
>         Attachments: FOP-3346-selector-drift.fo, fop-test-fonts.xconf
>
>
> CIDSubset.getChars builds a char\[] with StringBuilder.appendCodePoint, so a 
> supplementary-plane character occupies two slots. PDFToUnicodeCMap then 
> derives the character selector from the array position. Every selector after 
> the pair is therefore written one too high.
> Measured with the 2.11 command line, DejaVu Math TeX Gyre, the text "A𝐀BZ": 
> the content stream uses selectors 3 4 5 6, and the CMap says
> {noformat}
> <0003> <0041>  <0004> <d835dc00>  <0006> <0042>  <0007> <005a>
> {noformat}
> Selector 5, the B, has no entry, and selector 6, the Z, is published as B. 
> pdftotext extracts "A𝐀 B"; pdf.js and PDFium give the missing selector as 
> U+0005.
> PDFToUnicodeCMapTestCase.surrogatePairTest pins the drift: it expects the 
> entry after the pair at 0x63 to be 0x65.
> h3. Fix
> Build the CMap from one destination per selector (a String, so a surrogate 
> pair is one entry of length two) rather than from a positional char\[]. The 
> range logic then needs no surrogate special cases: an entry may join a 
> bfrange when it is one code point, and two entries are consecutive when their 
> code points are and their selectors share a 256 block. Expectations in 
> surrogatePairTest, surrogatePairRangeTest, surrogatePairsRangeTest and 
> rangeSizeSurrogateTest change accordingly; the last also used low surrogates 
> that ran past U+DFFF and now starts at U+DC00.
> After the fix the same file extracts as "A𝐀BZ" in pdftotext, mupdf, pdf.js 
> and PDFium.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to