Should I specify that urls.txt file as /user/root/urls/urls.txt so it
pulls it off the ndfs?

On Fri, 2005-12-16 at 21:39 +0100, Stefan Groschupf wrote:
>  >I would like to crawl a list of domains,
> > but I would like crawling limited to just those domains. When I first
> > played around with nutch in a localsetup I just set the following
> > property in nutch-site.xml:
> > <property>
> >   <name>urlfilter.prefix.file</name>
> >   <value>urls.txt</value>
> >   <description>Name of file on CLASSPATH containing url prefixes
> >   used by urlfilter-prefix (PrefixURLFilter) plugin.</description>
> > </property>
> > Can I do this in a mapred system?
> Sure, all plugins also works with map reduce.
> There is also a db url filter plugin contribution in the jira.
> You need to check what is better for your, needs if you only have a  
> few host than the file based would be enough.
> 
> > Also how does the fetching work, does
> > each new round of generate crawldb fetch the next "level" of urls?
> Somehow yes, but nutch use a important urls first algorithm called opic.
> > I'm
> > wondering the best way to put the crawl/index system on autopilot so
> > pages are crawled and updated regularly.
> A shell script, with a some regular expression matching of the nutch  
> tools outcome.
> 
> > Thanks again for you help Stefan.
> 
> Nop, this is how open source works, just help other newbies as well  
> if you know how to do it.
> 
> Stefan
> 

Reply via email to