[
https://issues.apache.org/jira/browse/TIKA-2347?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16264877#comment-16264877
]
Hudson commented on TIKA-2347:
------------------------------
SUCCESS: Integrated in Jenkins build Tika-trunk #1396 (See
[https://builds.apache.org/job/Tika-trunk/1396/])
Fix for TIKA-2347 Adds underline extraction from word documents (david:
[https://github.com/apache/tika/commit/d64a32c63f376b9e003a4512adfec05414d4dfe6])
* (edit)
tika-parsers/src/test/java/org/apache/tika/parser/microsoft/WordParserTest.java
* (edit)
tika-parsers/src/test/java/org/apache/tika/parser/microsoft/ooxml/OOXMLParserTest.java
* (edit)
tika-parsers/src/main/java/org/apache/tika/parser/microsoft/WordExtractor.java
* (edit)
tika-parsers/src/main/java/org/apache/tika/parser/microsoft/ooxml/XWPFWordExtractorDecorator.java
TIKA-2347 - Added extraction of <strike> element in DOCX files (david:
[https://github.com/apache/tika/commit/93cbed6df993ef01e59c55b86449b664e9052cae])
* (edit)
tika-parsers/src/test/java/org/apache/tika/parser/microsoft/ooxml/OOXMLParserTest.java
* (edit)
tika-parsers/src/main/java/org/apache/tika/parser/microsoft/ooxml/XWPFWordExtractorDecorator.java
* (edit)
tika-parsers/src/test/java/org/apache/tika/parser/microsoft/WordParserTest.java
* (edit) tika-parsers/src/test/resources/test-documents/testWORD_various.docx
* (edit) tika-parsers/src/test/resources/test-documents/testWORD_various.doc
TIKA-2347 - Add underline extraction from Word documents (doc/docx) from
(david:
[https://github.com/apache/tika/commit/639f3bf361a08210da8fae68e3eeb4e12df6c4de])
* (edit) CHANGES.txt
TIKA-2347 - Add underline extraction from Word documents (doc/docx) from
(david:
[https://github.com/apache/tika/commit/beedc4277526c6327524acb3a799b2ca6c898a05])
* (edit) CHANGES.txt
> Underlined text is not decorated as such when extracting from word documents
> ----------------------------------------------------------------------------
>
> Key: TIKA-2347
> URL: https://issues.apache.org/jira/browse/TIKA-2347
> Project: Tika
> Issue Type: Bug
> Components: parser
> Affects Versions: 2.0, 1.14
> Reporter: Stuart Hendren
> Assignee: Dave Meikle
> Fix For: 1.17
>
>
> When extracting from doc and docx bold and italic text decoration is
> extracted, however underlining is not. Can be demonstrated in WordParserTest
> or OOXMLParserTest (change to docx) with the following test case.
> {code:title=WordParserTest.java|borderStyle=solid}
> @Test
> public void testTextDecoration() throws Exception {
> XMLResult result = getXML("testWORD_various.doc");
> String xml = result.xml;
> assertTrue(xml.contains("<b>Bold</b>"));
> assertTrue(xml.contains("<i>italic</i>"));
> assertTrue(xml.contains("<u>underline</u>"));
> }
> {code}
--
This message was sent by Atlassian JIRA
(v6.4.14#64029)