[
https://issues.apache.org/jira/browse/NUTCH-961?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15170554#comment-15170554
]
ASF GitHub Bot commented on NUTCH-961:
--------------------------------------
Github user lewismc commented on a diff in the pull request:
https://github.com/apache/nutch/pull/92#discussion_r54332145
--- Diff: conf/nutch-default.xml ---
@@ -876,6 +876,19 @@
</description>
</property>
+<!-- tika properties -->
+
+<property>
+ <name>tika.boilerpipe</name>
+ <value>false</value>
--- End diff --
Can you provide descriptions of these properties please?
> Expose Tika's boilerpipe support
> --------------------------------
>
> Key: NUTCH-961
> URL: https://issues.apache.org/jira/browse/NUTCH-961
> Project: Nutch
> Issue Type: New Feature
> Components: parser
> Affects Versions: 1.11
> Reporter: Markus Jelsma
> Assignee: Markus Jelsma
> Fix For: 1.12
>
> Attachments: BoilerpipeExtractorRepository.java,
> NUTCH-961-1.11-1.patch, NUTCH-961-1.3-3.patch,
> NUTCH-961-1.3-tikaparser.patch, NUTCH-961-1.3-tikaparser1.patch,
> NUTCH-961-1.4-dombuilder-1.patch, NUTCH-961-1.5-1.patch,
> NUTCH-961-1.8-1.patch, NUTCH-961-2.1-v1.patch, NUTCH-961-2.1-v2.patch,
> NUTCH-961.patch, NUTCH-961.patch, NUTCH-961v2.patch,
> nutch-2.x-boilerpipe.patch
>
>
> Tika 0.8 comes with the Boilerpipe content handler which can be used to
> extract boilerplate content from HTML pages. We should see how we can expose
> Boilerplate in the Nutch cofiguration.
> Use the following properties to enable and control Boilerpipe.
> {code}
> <property>
> <name>tika.extractor</name>
> <value>none</value>
> <description>
> Which text extraction algorithm to use. Valid values are: boilerpipe or
> none.
> </description>
> </property>
>
> <property>
> <name>tika.extractor.boilerpipe.algorithm</name>
> <value>ArticleExtractor</value>
> <description>
> Which Boilerpipe algorithm to use. Valid values are: DefaultExtractor,
> ArticleExtractor
> or CanolaExtractor.
> </description>
> </property>
> {code}
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)