[ 
https://issues.apache.org/jira/browse/NUTCH-2457?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16940884#comment-16940884
 ] 

ASF GitHub Bot commented on NUTCH-2457:
---------------------------------------

sebastian-nagel commented on issue #474: NUTCH-2457 Embedded documents likely 
not correctly parsed by Tika
URL: https://github.com/apache/nutch/pull/474#issuecomment-536520812
 
 
   (rebased to master, resolved conflicts)
 
----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
[email protected]


> Embedded documents likely not correctly parsed by Tika
> ------------------------------------------------------
>
>                 Key: NUTCH-2457
>                 URL: https://issues.apache.org/jira/browse/NUTCH-2457
>             Project: Nutch
>          Issue Type: Bug
>          Components: parser, plugin
>    Affects Versions: 1.14
>            Reporter: Tim Allison
>            Priority: Major
>              Labels: patch-available
>             Fix For: 1.16
>
>
> While working on TIKA-2490, I think I found that Nutch's current method of 
> requesting a mime-specific parser for each file will fail to parse embedded 
> files, e.g. 
> https://github.com/apache/tika/blob/master/tika-server/src/test/resources/test_recursive_embedded.docx
> The fix should be straightforward, and I'll submit a PR once I can get Nutch 
> up and running in my dev environment. 



--
This message was sent by Atlassian Jira
(v8.3.4#803005)

Reply via email to