[ 
https://issues.apache.org/jira/browse/PDFBOX-6256?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119826#comment-18119826
 ] 

Tilman Hausherr commented on PDFBOX-6256:
-----------------------------------------

Re inline, see the file from PDFBOX-3248 and PDFBOX-5868 (or the files created 
after the change). I hadn't bothered to look at the content stream until now 
because the tests looked good and passed. But copilot is right. This is a 
similar problem to what we had with MCID in PDFBOX-5890. We could create 
another beginMarkedContent() method with a string, but I'm not really sure, and 
we can still do it later.

> Extracted Text is incorrect. Option for ActualText proposed.
> ------------------------------------------------------------
>
>                 Key: PDFBOX-6256
>                 URL: https://issues.apache.org/jira/browse/PDFBOX-6256
>             Project: PDFBox
>          Issue Type: Bug
>    Affects Versions: 3.0.8 PDFBox
>            Reporter: Volker Kunert
>            Priority: Major
>             Fix For: 3.0.9 PDFBox, 4.0.0
>
>
> For a correct result of text extraction an option to use ActualText is added 
> to GlyphLayoutProcessor(Awt|Fop)
> SeeĀ https://github.com/apache/pdfbox/pull/526



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to