[
https://issues.apache.org/jira/browse/NUTCH-2202?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15186822#comment-15186822
]
ASF GitHub Bot commented on NUTCH-2202:
---------------------------------------
GitHub user lewismc opened a pull request:
https://github.com/apache/nutch/pull/97
NUTCH-2202 Integration of Anthelion (Focused Crawling Module) into Nutch
This is a first pass at integration of the Anthelion plugin for Nutch.
@RobertMeusel, can you please scope this out?
The initial question I need answered is whether we need to ship Anthelion
itself with Nutch? This patch does exactly that.
You can merge this pull request into a Git repository by running:
$ git pull https://github.com/lewismc/nutch NUTCH-2202
Alternatively you can review and apply these changes as the patch at:
https://github.com/apache/nutch/pull/97.patch
To close this pull request, make a commit to your master/trunk branch
with (at least) the following in the commit message:
This closes #97
----
commit d994a2e7a1dca83e42800987a300ca7147d8608b
Author: Lewis John McGibbney <[email protected]>
Date: 2016-03-08T18:44:49Z
NUTCH-2202 Integration of Anthelion (Focused Crawling Module) into Nutch
----
> Integration of Anthelion (Focused Crawling Module) into Nutch
> -------------------------------------------------------------
>
> Key: NUTCH-2202
> URL: https://issues.apache.org/jira/browse/NUTCH-2202
> Project: Nutch
> Issue Type: Improvement
> Components: parser, scoring
> Reporter: Robert Meusel
> Assignee: Lewis John McGibbney
> Labels: any23, online_learning
>
> We have recently released anthelion, which is a focused crawler plugin for
> structured data which can be extracted with any23.
> (https://github.com/yahoo/anthelion) As proposed by Lewis (Lewis John
> McGibbney) we think the integration of the parser (any23) and the scoring
> function based on the online learner could be a good improvement for nutch.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)