I'm also guessing that it's important for all tasktrackers to have the appropriate configuration set in their conf/nutch-site.xml or can I just do it on the namenode?
On Fri, 2005-12-16 at 12:57 -0800, Michael Taggart wrote: > Should I specify that urls.txt file as /user/root/urls/urls.txt so it > pulls it off the ndfs? > > On Fri, 2005-12-16 at 21:39 +0100, Stefan Groschupf wrote: > > >I would like to crawl a list of domains, > > > but I would like crawling limited to just those domains. When I first > > > played around with nutch in a localsetup I just set the following > > > property in nutch-site.xml: > > > <property> > > > <name>urlfilter.prefix.file</name> > > > <value>urls.txt</value> > > > <description>Name of file on CLASSPATH containing url prefixes > > > used by urlfilter-prefix (PrefixURLFilter) plugin.</description> > > > </property> > > > Can I do this in a mapred system? > > Sure, all plugins also works with map reduce. > > There is also a db url filter plugin contribution in the jira. > > You need to check what is better for your, needs if you only have a > > few host than the file based would be enough. > > > > > Also how does the fetching work, does > > > each new round of generate crawldb fetch the next "level" of urls? > > Somehow yes, but nutch use a important urls first algorithm called opic. > > > I'm > > > wondering the best way to put the crawl/index system on autopilot so > > > pages are crawled and updated regularly. > > A shell script, with a some regular expression matching of the nutch > > tools outcome. > > > > > Thanks again for you help Stefan. > > > > Nop, this is how open source works, just help other newbies as well > > if you know how to do it. > > > > Stefan > >
