[jira] [Commented] (TIKA-3657) Microsoft documents are not text parsed when running under Docker

Tim Barrett (Jira) Tue, 25 Jan 2022 03:10:05 -0800


    [ 
https://issues.apache.org/jira/browse/TIKA-3657?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17481736#comment-17481736
 ]


Tim Barrett commented on TIKA-3657:
-----------------------------------

I managed to build our app as follows:

<dependency>

<groupId>nalanda.downloads</groupId>

 <artifactId>tika-core</artifactId>

<version>2.2.2-SNAPSHOT</version>

 </dependency>

 

 <dependency>

<groupId>nalanda.downloads</groupId>

<artifactId>tika-parsers-standard-package</artifactId>

<version>2.2.2-SNAPSHOT</version>

 

 <exclusions>

 <exclusion>

<groupId>xml-apis</groupId>

 <artifactId>xml-apis</artifactId>

 </exclusion>

 </exclusions>

 

 </dependency>

 

 <dependency>

<groupId>nalanda.downloads</groupId>

<artifactId>tika-parser-microsoft-module</artifactId>

<version>2.2.2-SNAPSHOT</version>

 </dependency>

 

 <dependency>

<groupId>nalanda.downloads</groupId>

<artifactId>tika-parser-pdf-module</artifactId>

<version>2.2.2-SNAPSHOT</version>

 </dependency>

 

 

<!-- https://mvnrepository.com/artifact/org.apache.pdfbox/pdfbox -->

 <dependency>

<groupId>org.apache.pdfbox</groupId>

 <artifactId>pdfbox</artifactId>

<version>2.0.25</version>

 </dependency>

 

<!-- https://mvnrepository.com/artifact/org.apache.pdfbox/pdfbox-tools -->

 <dependency>

<groupId>org.apache.pdfbox</groupId>

<artifactId>pdfbox-tools</artifactId>

<version>2.0.25</version>

 </dependency>

 

<!-- https://mvnrepository.com/artifact/org.apache.commons/commons-compress -->

 <dependency>

<groupId>org.apache.commons</groupId>

<artifactId>commons-compress</artifactId>

 <version>1.21</version>

 </dependency>

 

<!-- https://mvnrepository.com/artifact/org.apache.poi/poi-ooxml -->

 <dependency>

<groupId>org.apache.poi</groupId>

 <artifactId>poi-ooxml</artifactId>

 <version>5.2.0</version>

 <exclusions>

 <exclusion>

<groupId>xml-apis</groupId>

 <artifactId>xml-apis</artifactId>

 </exclusion>

 </exclusions>

 </dependency>

 

> Microsoft documents are not text parsed when running under Docker
> -----------------------------------------------------------------
>
>                 Key: TIKA-3657
>                 URL: https://issues.apache.org/jira/browse/TIKA-3657
>             Project: Tika
>          Issue Type: Bug
>          Components: config, core, depedency
>    Affects Versions: 2.2.0, 2.2.1
>            Reporter: Tim Barrett
>            Priority: Major
>             Fix For: 2.2.2
>
>         Attachments: tika-config.xml
>
>
> We use EmbeddedDocumentExtractor, with this code:
> NalyticsEmbeddedDocumentExtractor nalyticsEmbeddedDocumentExtractor = *new* 
> NalyticsEmbeddedDocumentExtractor(*this*);
> *this*.context.set(EmbeddedDocumentExtractor.*class*, 
> nalyticsEmbeddedDocumentExtractor);
> This all works fine for us, and has been used in production for a few years. 
> This also works under Tika 2.2.0 when running in development environments 
> (Eclipse, Apache Tomcat). However when running under Docker the text 
> withinMicrosoft documents (Word etc) is not parsed. Under Tika 2.1.0, under 
> Docker, the Microsoft documents are fully parsed, so this problem was 
> introduced in 2.2.0
> Interestingly, I found that if *anything at all* is added to the context via 
> context.set the same problem occurs. Also, if the standard Tika Embedded 
> Document Extractor is used the same problem occurs. Our Docker image contains 
> our application's code which uses Tika, as well as Apache DS. The problem 
> occurs running Docker on Ubuntu, Mac OS and Windows.
>  



--
This message was sent by Atlassian Jira
(v8.20.1#820001)

[jira] [Commented] (TIKA-3657) Microsoft documents are not text parsed when running under Docker

Reply via email to