[
https://issues.apache.org/jira/browse/NUTCH-961?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15148758#comment-15148758
]
Hudson commented on NUTCH-961:
------------------------------
SUCCESS: Integrated in Nutch-trunk #3347 (See
[https://builds.apache.org/job/Nutch-trunk/3347/])
NUTCH-961 Expose Tika's Boilerpipe support (markus:
[http://svn.apache.org/viewvc/nutch/trunk/?view=rev&rev=1730694])
* trunk/CHANGES.txt
* trunk/conf/nutch-default.xml
*
trunk/src/plugin/parse-tika/src/java/org/apache/nutch/parse/tika/BoilerpipeExtractorRepository.java
*
trunk/src/plugin/parse-tika/src/java/org/apache/nutch/parse/tika/TikaParser.java
> Expose Tika's boilerpipe support
> --------------------------------
>
> Key: NUTCH-961
> URL: https://issues.apache.org/jira/browse/NUTCH-961
> Project: Nutch
> Issue Type: New Feature
> Components: parser
> Affects Versions: 1.11
> Reporter: Markus Jelsma
> Assignee: Markus Jelsma
> Fix For: 1.12
>
> Attachments: BoilerpipeExtractorRepository.java,
> NUTCH-961-1.11-1.patch, NUTCH-961-1.3-3.patch,
> NUTCH-961-1.3-tikaparser.patch, NUTCH-961-1.3-tikaparser1.patch,
> NUTCH-961-1.4-dombuilder-1.patch, NUTCH-961-1.5-1.patch,
> NUTCH-961-1.8-1.patch, NUTCH-961-2.1-v1.patch, NUTCH-961-2.1-v2.patch,
> NUTCH-961.patch, NUTCH-961.patch, NUTCH-961v2.patch,
> nutch-2.x-boilerpipe.patch
>
>
> Tika 0.8 comes with the Boilerpipe content handler which can be used to
> extract boilerplate content from HTML pages. We should see how we can expose
> Boilerplate in the Nutch cofiguration.
> Use the following properties to enable and control Boilerpipe.
> {code}
> <property>
> <name>tika.extractor</name>
> <value>none</value>
> <description>
> Which text extraction algorithm to use. Valid values are: boilerpipe or
> none.
> </description>
> </property>
>
> <property>
> <name>tika.extractor.boilerpipe.algorithm</name>
> <value>ArticleExtractor</value>
> <description>
> Which Boilerpipe algorithm to use. Valid values are: DefaultExtractor,
> ArticleExtractor
> or CanolaExtractor.
> </description>
> </property>
> {code}
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)