[ 
https://issues.apache.org/jira/browse/PDFBOX-5838?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17863477#comment-17863477
 ] 

Andreas Lehmkühler commented on PDFBOX-5838:
--------------------------------------------

Sorry for the late answer. I had another look at the results of the regression 
tests. There were more files which got worse, BUT at least the files I'm able 
to read, were already bad in the first place. The changes didn't really improve 
the extraction results but simply replace one bad result with another.

Saying that, I stick to my intention to keep the current implementation and 
concur with [~tilman] to close this ticket.

> Text extraction garbled in this file, was OK in 3.0.2 / 2.0.31
> --------------------------------------------------------------
>
>                 Key: PDFBOX-5838
>                 URL: https://issues.apache.org/jira/browse/PDFBOX-5838
>             Project: PDFBox
>          Issue Type: Bug
>          Components: Text extraction
>    Affects Versions: 2.0.32, 3.0.3 PDFBox
>            Reporter: Tilman Hausherr
>            Priority: Major
>              Labels: regression
>         Attachments: OFLSV3YFD3TDOU4YZTL2QY745W53W3DW.pdf, 
> PDFBOX-5838-0024320-reduced.pdf
>
>
> discovered in 2.0.32 regression tests



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to