You can set the property "Protocol.CHECK_ROBOTS" to false in nutch-site.xml to ignore robots.txt. - Sathyam
Vijay Krishnan <[EMAIL PROTECTED]> wrote: Hi all, Do you know what file in Nutch parses robots.txt? If there is no option in Nutch to ignore robots.txt, I would at least like to modify the source. Thanks, Vijay 2008/5/25 Ivannie : > Vijay Krishnan, > > if the site does not allow you to crawl, why not choose another? > > otherwise you can check the code of parsing robots.txt and disable it > > > >>Hi all, >> >> I wish to crawl a certain set of URLs to depth 1 (without any >>deeper crawling) for further analysis. I find that nutch does not >>crawl URLs which do not have the requisite permissions in robots.txt. >>Is there some way I can disable nutch from looking at robots.txt? That >>will make my job much easier than trying to save the webpages some >>other way and then passing it through nutch. >> >> >>Thanks >>-- >>Vijay Krishnan >>http://www.cs.stanford.edu/~vijayk > > = = = = = = = = = = = = = = = = = = = = > > > > > Ivannie > [EMAIL PROTECTED] > 2008-05-26 > >
