Vijay Krishnan,
if the site does not allow you to crawl, why not choose another?
otherwise you can check the code of parsing robots.txt and disable it
>Hi all,
>
> I wish to crawl a certain set of URLs to depth 1 (without any
>deeper crawling) for further analysis. I find that nutch does not
>crawl URLs which do not have the requisite permissions in robots.txt.
>Is there some way I can disable nutch from looking at robots.txt? That
>will make my job much easier than trying to save the webpages some
>other way and then passing it through nutch.
>
>
>Thanks
>--
>Vijay Krishnan
>http://www.cs.stanford.edu/~vijayk
= = = = = = = = = = = = = = = = = = = =
Ivannie
[EMAIL PROTECTED]
2008-05-26