[ 
https://issues.apache.org/jira/browse/NUTCH-61?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#action_12464725
 ] 

Armel Nene commented on NUTCH-61:
---------------------------------

I was able to apply the patch to Nutch 0.8.1 and have it successfully running. 
I think this patch should be part of the core code. When crawling a terrabyte 
of data, it is important that only changed data be fetched and parsed. Prior to 
apply this patch, we run Nutch in our lab and were confronted with SYSTEM OUT 
MEMORY messages when trying to crawl files as small as 10Gb of data. Now with 
this patch, it's true the performance will be slower because of checking for 
the unmodified data but overall it's worth it.

+5 for this patch.



> Adaptive re-fetch interval. Detecting umodified content
> -------------------------------------------------------
>
>                 Key: NUTCH-61
>                 URL: https://issues.apache.org/jira/browse/NUTCH-61
>             Project: Nutch
>          Issue Type: New Feature
>          Components: fetcher
>            Reporter: Andrzej Bialecki 
>         Assigned To: Andrzej Bialecki 
>         Attachments: 20050606.diff, 20051230.txt, 20060227.txt, 
> nutch-61-417287.patch
>
>
> Currently Nutch doesn't adjust automatically its re-fetch period, no matter 
> if individual pages change seldom or frequently. The goal of these changes is 
> to extend the current codebase to support various possible adjustments to 
> re-fetch times and intervals, and specifically a re-fetch schedule which 
> tries to adapt the period between consecutive fetches to the period of 
> content changes.
> Also, these patches implement checking if the content has changed since last 
> fetching; protocol plugins are also changed to make use of this information, 
> so that if content is unmodified it doesn't have to be fetched and processed.

-- 
This message is automatically generated by JIRA.
-
If you think it was sent incorrectly contact one of the administrators: 
https://issues.apache.org/jira/secure/Administrators.jspa
-
For more information on JIRA, see: http://www.atlassian.com/software/jira

        

-------------------------------------------------------------------------
Take Surveys. Earn Cash. Influence the Future of IT
Join SourceForge.net's Techsay panel and you'll get the chance to share your
opinions on IT & business topics through brief surveys - and earn cash
http://www.techsay.com/default.php?page=join.php&p=sourceforge&CID=DEVDEV
_______________________________________________
Nutch-developers mailing list
Nutch-developers@lists.sourceforge.net
https://lists.sourceforge.net/lists/listinfo/nutch-developers

Reply via email to