You can set the property "Protocol.CHECK_ROBOTS" to false in nutch-site.xml to 
ignore robots.txt. 
   
  - Sathyam
  

Vijay Krishnan <[EMAIL PROTECTED]> wrote:
  Hi all,

Do you know what file in Nutch parses robots.txt? If there is no
option in Nutch to ignore robots.txt, I would at least like to modify
the source.


Thanks,
Vijay

2008/5/25 Ivannie :
> Vijay Krishnan,
>
> if the site does not allow you to crawl, why not choose another?
>
> otherwise you can check the code of parsing robots.txt and disable it
>
>
>
>>Hi all,
>>
>> I wish to crawl a certain set of URLs to depth 1 (without any
>>deeper crawling) for further analysis. I find that nutch does not
>>crawl URLs which do not have the requisite permissions in robots.txt.
>>Is there some way I can disable nutch from looking at robots.txt? That
>>will make my job much easier than trying to save the webpages some
>>other way and then passing it through nutch.
>>
>>
>>Thanks
>>--
>>Vijay Krishnan
>>http://www.cs.stanford.edu/~vijayk
>
> = = = = = = = = = = = = = = = = = = = =
>
>
>
>
> Ivannie
> [EMAIL PROTECTED]
> 2008-05-26
>
>


       

Reply via email to