[jira] [Commented] (TIKA-1699) Integrate the GROBID PDF extractor in Tika

Chris A. Mattmann (JIRA) Sun, 16 Aug 2015 21:10:15 -0700

    [ 
https://issues.apache.org/jira/browse/TIKA-1699?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14698988#comment-14698988
 ]


Chris A. Mattmann commented on TIKA-1699:
-----------------------------------------

OK I got the fully REST services version of the GROBID PDF parser implemented. 
Tests are passing and I'm going to commit it within the next few minutes. 
Basically it only adds the CXF rest client dependency and also the org.json 
dependency. Lot better, and lot smaller. Also GROBID can exist on another 
machine now. Will update the docs shortly.

> Integrate the GROBID PDF extractor in Tika
> ------------------------------------------
>
>                 Key: TIKA-1699
>                 URL: https://issues.apache.org/jira/browse/TIKA-1699
>             Project: Tika
>          Issue Type: New Feature
>          Components: parser
>            Reporter: Sujen Shah
>            Assignee: Chris A. Mattmann
>              Labels: memex
>             Fix For: 1.11
>
>         Attachments: TIKA-1699.grobid-core.MattmannShah.081515.patch.txt, 
> TIKA-1699.restgrobid.MattmannWIP081515.patch.txt
>
>
> GROBID (http://grobid.readthedocs.org/en/latest/) is a machine learning 
> library for extracting, parsing and re-structuring raw documents such as PDF 
> into structured TEI-encoded documents with a particular focus on technical 
> and scientific publications.
> It has a java api which can be used to augment PDF parsing for journals and 
> help extract extra metadata about the paper like authors, publication, 
> citations, etc. 
> It would be nice to have this integrated into Tika, I have tried it on my 
> local, will issue a pull request soon.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

[jira] [Commented] (TIKA-1699) Integrate the GROBID PDF extractor in Tika

Reply via email to