The current patch actually wouldn't allow the crawling of the outlinks. If a page is filtered the content and parse objects wouldn't be saved and the URL would be poisoned in the CrawlDb from being crawled in the future. Maybe we can add in the ability to still crawl outlinks. I think this would need to be a global setting applied to all pages filtered versus a page by page basis.

It is a new extension point so other plugins can be written, such as one interacting with Mahout algorithms.

Dennis

On 06/09/2010 01:14 AM, David Stuart wrote:
I haven't review this patch but I was about to start work on something similar so a big +1 on the ability to filter out the page but allow crawling of the outlinks. Also if the filter was able to be pluggable to external inputs (like mahout) def +1 on that too

David Stuart

On 9 Jun 2010, at 06:53, "Andrzej Bialecki (JIRA)" <[email protected]> wrote:


[ https://issues.apache.org/jira/browse/NUTCH-828?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12876964#action_12876964 ]

Andrzej Bialecki  commented on NUTCH-828:
-----------------------------------------

First, as you point out, we cannot ignore the page because the problem will repeat itself as we keep re-discovering it, so we have to "poison" it with GONE - and I think it's ok to add another status here to express that we never ever want to collect this page, because GONE gets reset periodically.

If we run Fetcher in parsing mode then we can change this status immediately, so no problem here. If we run ParseSegment then we can also update this status in a similar way as we implement the signature update, i.e. in ParseOutputFormat emit a <pageUrl,CrawlDatum> that will switch the status of this page when collected later on in CrawlDbReducer.

Fetch Filter
------------

               Key: NUTCH-828
               URL: https://issues.apache.org/jira/browse/NUTCH-828
           Project: Nutch
        Issue Type: New Feature
        Components: fetcher
       Environment: All
          Reporter: Dennis Kubes
          Assignee: Dennis Kubes
           Fix For: 1.1

Attachments: NUTCH-828-1-20100608.patch, NUTCH-828-2-20100608.patch


Adds a Nutch extension point for a fetch filter. The fetch filter allows filtering content and parse data/text after it is fetched but before it is written to segments. The fliter can return true if content is to be written or false if it is not. Some use cases for this filter would be topical search engines that only want to fetch/index certain types of content, for example a news or sports only search engine. In these types of situations the only way to determine if content belongs to a particular set is to fetch the page and then analyze the content. If the content passes, meaning belongs to the set of say sports pages, then we want to include it. If it doesn't then we want to ignore it, never fetch that same page in the future, and ignore any urls on that page. If content is rejected due to a fetch filter then its status is written to the CrawlDb as gone and its content is ignored and not written to segments. This effectively stop crawling along the crawl path of that page and the urls from that page. An example filter, fetch-safe, is provided that allows fetching content that does not contain a list of bad words.

--
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.

Reply via email to