[
https://issues.apache.org/jira/browse/NUTCH-961?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15170557#comment-15170557
]
ASF GitHub Bot commented on NUTCH-961:
--------------------------------------
Github user lewismc commented on a diff in the pull request:
https://github.com/apache/nutch/pull/92#discussion_r54332201
--- Diff:
src/plugin/parse-tika/src/java/org/apache/nutch/parse/tika/TikaParser.java ---
@@ -109,7 +114,18 @@ public Parse getParse(String url, WebPage page) {
HTMLDocumentImpl doc = new HTMLDocumentImpl();
doc.setErrorChecking(false);
DocumentFragment root = doc.createDocumentFragment();
- DOMBuilder domhandler = new DOMBuilder(doc, root);
+ // DOMBuilder domhandler = new DOMBuilder(doc, root);
+ ContentHandler domHandler;
+ // Check whether to use Tika's BoilerplateContentHandler
+ if (useBoilerpipe) {
+ LOG.debug("Using Tikas's Boilerpipe with Extractor: " +
boilerpipeExtractorName);
--- End diff --
Can also use more efficient slf4j convention
logger.debug("The entry is {}.", entry);
> Expose Tika's boilerpipe support
> --------------------------------
>
> Key: NUTCH-961
> URL: https://issues.apache.org/jira/browse/NUTCH-961
> Project: Nutch
> Issue Type: New Feature
> Components: parser
> Affects Versions: 1.11
> Reporter: Markus Jelsma
> Assignee: Markus Jelsma
> Fix For: 1.12
>
> Attachments: BoilerpipeExtractorRepository.java,
> NUTCH-961-1.11-1.patch, NUTCH-961-1.3-3.patch,
> NUTCH-961-1.3-tikaparser.patch, NUTCH-961-1.3-tikaparser1.patch,
> NUTCH-961-1.4-dombuilder-1.patch, NUTCH-961-1.5-1.patch,
> NUTCH-961-1.8-1.patch, NUTCH-961-2.1-v1.patch, NUTCH-961-2.1-v2.patch,
> NUTCH-961.patch, NUTCH-961.patch, NUTCH-961v2.patch,
> nutch-2.x-boilerpipe.patch
>
>
> Tika 0.8 comes with the Boilerpipe content handler which can be used to
> extract boilerplate content from HTML pages. We should see how we can expose
> Boilerplate in the Nutch cofiguration.
> Use the following properties to enable and control Boilerpipe.
> {code}
> <property>
> <name>tika.extractor</name>
> <value>none</value>
> <description>
> Which text extraction algorithm to use. Valid values are: boilerpipe or
> none.
> </description>
> </property>
>
> <property>
> <name>tika.extractor.boilerpipe.algorithm</name>
> <value>ArticleExtractor</value>
> <description>
> Which Boilerpipe algorithm to use. Valid values are: DefaultExtractor,
> ArticleExtractor
> or CanolaExtractor.
> </description>
> </property>
> {code}
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)