Hi all,
Do you know what file in Nutch parses robots.txt? If there is no
option in Nutch to ignore robots.txt, I would at least like to modify
the source.
Thanks,
Vijay
2008/5/25 Ivannie <[EMAIL PROTECTED]>:
> Vijay Krishnan,
>
> if the site does not allow you to crawl, why not choose another?
>
> otherwise you can check the code of parsing robots.txt and disable it
>
>
>
>>Hi all,
>>
>> I wish to crawl a certain set of URLs to depth 1 (without any
>>deeper crawling) for further analysis. I find that nutch does not
>>crawl URLs which do not have the requisite permissions in robots.txt.
>>Is there some way I can disable nutch from looking at robots.txt? That
>>will make my job much easier than trying to save the webpages some
>>other way and then passing it through nutch.
>>
>>
>>Thanks
>>--
>>Vijay Krishnan
>>http://www.cs.stanford.edu/~vijayk
>
> = = = = = = = = = = = = = = = = = = = =
>
>
>
>
> Ivannie
> [EMAIL PROTECTED]
> 2008-05-26
>
>