Awesome!  That fixes it indeed.
Great work again! 
 
-----Original message-----
> From:Tejas Patil <[email protected]>
> Sent: Friday 24th January 2014 17:38
> To: [email protected]
> Subject: Re: Order of robots file
> 
> Hi Markus,
> 
> I found the root cause :) In RobotRulesParser, when Nutch creates a agent
> string from multiple agents, it combines agents from both 'http.agent.name'
> and 'http.robots.agents'. Along with that it appends a "*" to it in the
> end. This is sent to crawler commons while parsing the rules. eg. For the
> config you shared, the effective robots agent string formed by Nutch was
> "Mozilla, nutch, *". The "*" added to the end got matched with the first
> rule in robots file and thus resulted in the url being robots denied.
> 
> This is effectively:
> $ bin/nutch org.apache.nutch.protocol.RobotRulesParser robots.txt urls.txt
>  Mozilla,nutch,*
> not allowed: http://temp.com/index.html
> 
> Commenting out line 125 in
> src/java/org/apache/nutch/protocol/RobotRulesParser.java fixes the issue.
> 125      combinedAgentsString.append(", *");
> 
> Can you confirm the same ?
> 
> Thanks,
> Tejas
> 
> 
> On Fri, Jan 24, 2014 at 9:07 PM, Markus Jelsma
> <[email protected]>wrote:
> 
> > Thanks for confirming, now we at least know my checkout isn't funky :)
> >
> > -----Original message-----
> > > From:Tejas Patil <[email protected]>
> > > Sent: Friday 24th January 2014 16:34
> > > To: [email protected]
> > > Subject: Re: Order of robots file
> > >
> > > It got robots denied when I ran a full crawl (ie. inject -> generate ..
> > > etc). Strange why RobotRulesParser gave opposite result. I am looking
> > into
> > > this.
> > >
> > > ~tejas
> > >
> > >
> > > On Fri, Jan 24, 2014 at 8:48 PM, Markus Jelsma
> > > <[email protected]>wrote:
> > >
> > > > Can you try doing a simple crawl (freegen, fetch, parse, update) over 1
> > > > record?
> > > >
> > > > This is my relevant config:
> > > >
> > > >   <property>
> > > >     <name>http.agent.name</name>
> > > >     <value>Mozilla</value>
> > > >   </property>
> > > >   <property>
> > > >     <name>http.agent.version</name>
> > > >     <value>5.0</value>
> > > >   </property>
> > > >   <property>
> > > >     <name>http.robots.agents</name>
> > > >     <value>nutch</value>
> > > >   </property>
> > > >   <property>
> > > >     <name>http.agent.description</name>
> > > >     <value>compatible; NutchCrawler</value>
> > > >   </property>
> > > >   <property>
> > > >     <name>http.agent.url</name>
> > > >     <value>+http://example.org/</value>
> > > >   </property>
> > > >
> > > > # cat /var/www/robots.txt
> > > > User-agent: *
> > > > Disallow: /
> > > >
> > > > User-agent: nutch
> > > > Allow: /
> > > >
> > > > markus@midas:~/projects/apache/nutch/trunk/runtime/local$ rm -r
> > > > crawl/crawldb/ ; ./run.sh ; bin/nutch readdb crawl/crawldb/ -url
> > > > http://localhost/
> > > > FreeGenerator: starting at 2014-01-24 15:58:38
> > > > FreeGenerator: finished at 2014-01-24 15:58:40, elapsed: 00:00:02
> > > > Fetcher: Your 'http.agent.name' value should be listed first in
> > > > 'http.robots.agents' property.
> > > > Fetcher: starting at 2014-01-24 15:58:41
> > > > Fetcher: segment: crawl/segments/20140124155839
> > > > Fetcher Timelimit set for : 1390578521635
> > > > Using queue mode : byHost
> > > > Fetcher: threads: 10
> > > > Fetcher: time-out divisor: 4
> > > > QueueFeeder finished: total 1 records + hit by time limit :0
> > > > Using queue mode : byHost
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > fetching http://localhost/ (queue crawl delay=5000ms)
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Fetcher: throughput threshold: 1
> > > > Fetcher: throughput threshold retries: 20
> > > > -finishing thread FetcherThread, activeThreads=0
> > > > -activeThreads=0, spinWaiting=0, fetchQueues.totalSize=0
> > > > -activeThreads=0
> > > > Fetcher: finished at 2014-01-24 15:58:44, elapsed: 00:00:02
> > > > CrawlDb update: starting at 2014-01-24 15:58:45
> > > > CrawlDb update: db: crawl/crawldb
> > > > CrawlDb update: segments: [crawl/segments/20140124155839]
> > > > CrawlDb update: additions allowed: true
> > > > CrawlDb update: URL normalizing: false
> > > > CrawlDb update: URL filtering: false
> > > > CrawlDb update: 404 purging: false
> > > > CrawlDb update: Merging segment data into db.
> > > > CrawlDb update: finished at 2014-01-24 15:58:46, elapsed: 00:00:01
> > > > URL: http://localhost/
> > > > Version: 7
> > > > Status: 3 (db_gone)
> > > > Fetch time: Mon Mar 10 15:58:42 CET 2014
> > > > Modified time: Thu Jan 01 01:00:00 CET 1970
> > > > Retries since fetch: 0
> > > > Retry interval: 3888000 seconds (45 days)
> > > > Score: 0.0
> > > > Signature: null
> > > > Metadata:
> > > >         _pst_=robots_denied(18), lastModified=0
> > > >
> > > >
> > > >
> > > >
> > > > Now, when i try again with robots reordered:
> > > >
> > > > # cat /var/www/robots.txt
> > > > User-agent: nutch
> > > > Allow: /
> > > >
> > > > User-agent: *
> > > > Disallow: /
> > > >
> > > > markus@midas:~/projects/apache/nutch/trunk/runtime/local$ rm -r
> > > > crawl/crawldb/ ; ./run.sh ; bin/nutch readdb crawl/crawldb/ -url
> > > > http://localhost/
> > > > FreeGenerator: starting at 2014-01-24 16:17:41
> > > > FreeGenerator: finished at 2014-01-24 16:17:44, elapsed: 00:00:02
> > > > Fetcher: Your 'http.agent.name' value should be listed first in
> > > > 'http.robots.agents' property.
> > > > Fetcher: starting at 2014-01-24 16:17:45
> > > > Fetcher: segment: crawl/segments/20140124161743
> > > > Fetcher Timelimit set for : 1390579665184
> > > > Using queue mode : byHost
> > > > Fetcher: threads: 10
> > > > Fetcher: time-out divisor: 4
> > > > QueueFeeder finished: total 1 records + hit by time limit :0
> > > > Using queue mode : byHost
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > fetching http://localhost/ (queue crawl delay=5000ms)
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Using queue mode : byHost
> > > > -finishing thread FetcherThread, activeThreads=1
> > > > Fetcher: throughput threshold: 1
> > > > Fetcher: throughput threshold retries: 20
> > > > -finishing thread FetcherThread, activeThreads=0
> > > > -activeThreads=0, spinWaiting=0, fetchQueues.totalSize=0
> > > > -activeThreads=0
> > > > Fetcher: finished at 2014-01-24 16:17:47, elapsed: 00:00:02
> > > > CrawlDb update: starting at 2014-01-24 16:17:48
> > > > CrawlDb update: db: crawl/crawldb
> > > > CrawlDb update: segments: [crawl/segments/20140124161743]
> > > > CrawlDb update: additions allowed: true
> > > > CrawlDb update: URL normalizing: false
> > > > CrawlDb update: URL filtering: false
> > > > CrawlDb update: 404 purging: false
> > > > CrawlDb update: Merging segment data into db.
> > > > CrawlDb update: finished at 2014-01-24 16:17:49, elapsed: 00:00:01
> > > > URL: http://localhost/
> > > > Version: 7
> > > > Status: 2 (db_fetched)
> > > > Fetch time: Sun Feb 23 16:17:46 CET 2014
> > > > Modified time: Fri Jan 24 16:17:46 CET 2014
> > > > Retries since fetch: 0
> > > > Retry interval: 2592000 seconds (30 days)
> > > > Score: 0.0
> > > > Signature: 49d1ab51eb6b764c6aec12bcb17ebba0
> > > > Metadata:
> > > >         Content-Type=text/html
> > > >         _pst_=success(1), lastModified=0
> > > >         _rs_=2
> > > >
> > > >
> > > >
> > > > I tried putting this also in a unit test but that did not work, perhaps
> > > > the problem is not only in one component?
> > > >
> > > > Thanks
> > > > Markus
> > > >
> > > >
> > > >
> > > > -----Original message-----
> > > > > From:Tejas Patil <[email protected]>
> > > > > Sent: Friday 24th January 2014 16:06
> > > > > To: [email protected]
> > > > > Subject: Re: Order of robots file
> > > > >
> > > > > With the current Nutch trunk, the scenario you described gets
> > reproduced
> > > > > with multiple agents but not with single agent:
> > > > >
> > > > > $ cat robots.txt
> > > > > User-agent: *
> > > > > Disallow: /
> > > > >
> > > > > User-agent: nutch
> > > > > Allow: /
> > > > >
> > > > > $ bin/nutch org.apache.nutch.protocol.RobotRulesParser robots.txt
> > > > urls.txt
> > > > > nutch,*
> > > > > not allowed: http://temp.com/index.html
> > > > >
> > > > > $ bin/nutch org.apache.nutch.protocol.RobotRulesParser robots.txt
> > > > urls.txt
> > > > > nutch
> > > > > allowed: http://temp.com/index.html
> > > > >
> > > > > You need not add additional "*" in the http.robots.agents string as
> > CC
> > > > > would internally add it for you and smartly do a match with
> > (User-agent:
> > > > *)
> > > > > when no agent names are matching with the rules from robots.txt
> > > > >
> > > > > Multiple agents without any "*"
> > > > > --------------------------------------------
> > > > > As CC is greedy, it matched "nutch" and allowed in the first two
> > examples
> > > > > below. In the last case, it correctly did mismatch.
> > > > >
> > > > > $ bin/nutch org.apache.nutch.protocol.RobotRulesParser robots.txt
> > > > urls.txt
> > > > > foobar,nutch
> > > > > allowed: http://temp.com/index.html
> > > > > $ bin/nutch org.apache.nutch.protocol.RobotRulesParser robots.txt
> > > > urls.txt
> > > > > nutch,foobar
> > > > > allowed: http://temp.com/index.html
> > > > > $ bin/nutch org.apache.nutch.protocol.RobotRulesParser robots.txt
> > > > urls.txt
> > > > > foobar
> > > > > not allowed: http://temp.com/index.html
> > > > >
> > > > > Multiple agents with "*"
> > > > > -----------------------------------
> > > > > Bad result as it would match with "*" as soon as it scans the first
> > rule
> > > > in
> > > > > the robots file.
> > > > >
> > > > > $ bin/nutch org.apache.nutch.protocol.RobotRulesParser robots.txt
> > > > urls.txt
> > > > > foobar,nutch,*
> > > > > not allowed: http://temp.com/index.html
> > > > > $ bin/nutch org.apache.nutch.protocol.RobotRulesParser robots.txt
> > > > urls.txt
> > > > > nutch,foobar,*
> > > > > not allowed: http://temp.com/index.html
> > > > >
> > > > > Thanks,
> > > > > Tejas
> > > > >
> > > > >
> > > > > On Fri, Jan 24, 2014 at 8:27 PM, Markus Jelsma
> > > > > <[email protected]>wrote:
> > > > >
> > > > > >
> > > > > > -----Original message-----
> > > > > > > From:Tejas Patil <[email protected]>
> > > > > > > Sent: Friday 24th January 2014 15:49
> > > > > > > To: [email protected]
> > > > > > > Subject: Re: Order of robots file
> > > > > > >
> > > > > > > I am working on the scenario you just pointed out. By Apache
> > Nutch,
> > > > you
> > > > > > > mean the current codebase with CC or version before that ?
> > > > > >
> > > > > > No, this is current trunk.
> > > > > >
> > > > > > >
> > > > > > > CC differs from original nutch code as CC has kinda a greedy
> > approach
> > > > > > > wherein it tries to get a match / mismatch after every line is
> > sees
> > > > from
> > > > > > > the robots file. While the time I was working on delegation of
> > robots
> > > > > > > parsing to Crawler commons (CC), I remember that there was
> > > > difference in
> > > > > > > the semantics of original parsing code and CC's implementation
> > for
> > > > > > multiple
> > > > > > > robots agents.
> > > > > > > Here was my observation at that time:
> > > > > > >
> > > > > >
> > > >
> > https://issues.apache.org/jira/browse/NUTCH-1031?focusedCommentId=13558217&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13558217
> > > > > >
> > > > > > The problem i describe persists with one and multiple robot agent
> > > > names.
> > > > > > So it is also a problem with:
> > > > > >   <property>
> > > > > >     <name>http.robots.agents</name>
> > > > > >     <value>nutch</value>
> > > > > >   </property>
> > > > > >
> > > > > >
> > > > > > >
> > > > > > > ~tejas
> > > > > > >
> > > > > > >
> > > > > > > On Fri, Jan 24, 2014 at 8:11 PM, Markus Jelsma
> > > > > > > <[email protected]>wrote:
> > > > > > >
> > > > > > > > Tejas, the problem exists in Apache Nutch as well. We'll take
> > > > > > localhost as
> > > > > > > > example and the following config and robots.txt
> > > > > > > >
> > > > > > > >
> > > > > > > > # cat /var/www/robots.txt
> > > > > > > > User-agent: *
> > > > > > > > Disallow: /
> > > > > > > >
> > > > > > > > User-agent: nutch
> > > > > > > > Allow: /
> > > > > > > >
> > > > > > > >
> > > > > > > >
> > > > > > > > config:
> > > > > > > >   <property>
> > > > > > > >     <name>http.agent.name</name>
> > > > > > > >     <value>Mozilla</value>
> > > > > > > >   </property>
> > > > > > > >   <property>
> > > > > > > >     <name>http.agent.version</name>
> > > > > > > >     <value>5.0</value>
> > > > > > > >   </property>
> > > > > > > >   <property>
> > > > > > > >     <name>http.robots.agents</name>
> > > > > > > >     <value>nutch,*</value>
> > > > > > > >   </property>
> > > > > > > >   <property>
> > > > > > > >     <name>http.agent.description</name>
> > > > > > > >     <value>compatible; NutchCrawler</value>
> > > > > > > >   </property>
> > > > > > > >   <property>
> > > > > > > >     <name>http.agent.url</name>
> > > > > > > >     <value>+http://example.org/</value>
> > > > > > > >   </property>
> > > > > > > >
> > > > > > > >
> > > > > > > > URL: http://localhost/
> > > > > > > > Version: 7
> > > > > > > > Status: 3 (db_gone)
> > > > > > > > Fetch time: Mon Mar 10 15:36:48 CET 2014
> > > > > > > > Modified time: Thu Jan 01 01:00:00 CET 1970
> > > > > > > > Retries since fetch: 0
> > > > > > > > Retry interval: 3888000 seconds (45 days)
> > > > > > > > Score: 0.0
> > > > > > > > Signature: null
> > > > > > > > Metadata:
> > > > > > > >         _pst_=robots_denied(18), lastModified=0
> > > > > > > >
> > > > > > > >
> > > > > > > >
> > > > > > > > Can you confirm?
> > > > > > > >
> > > > > > > >
> > > > > > > >
> > > > > > > >
> > > > > > > >
> > > > > > > > -----Original message-----
> > > > > > > > > From:Markus Jelsma <[email protected]>
> > > > > > > > > Sent: Friday 24th January 2014 15:29
> > > > > > > > > To: [email protected]
> > > > > > > > > Subject: RE: Order of robots file
> > > > > > > > >
> > > > > > > > > Hi, sorry for being unclear. You understand correctly, i had
> > to
> > > > > > change
> > > > > > > > the robots.txt order and put our crawler ABOVE User-Agent: *.
> > > > > > > > >
> > > > > > > > > I have tried a unit test for lib-http to demonstrate the
> > problem
> > > > but
> > > > > > i
> > > > > > > > does not fail. I think i did something wrong in merging the
> > code
> > > > base.
> > > > > > I'll
> > > > > > > > look further.
> > > > > > > > >
> > > > > > > > > markus@midas:~/projects/apache/nutch/trunk$ svn diff
> > > > > > > >
> > > > > >
> > > >
> > src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
> > > > > > > > > Index:
> > > > > > > >
> > > > > >
> > > >
> > src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
> > > > > > > > >
> > > > ===================================================================
> > > > > > > > > ---
> > > > > > > >
> > > > > >
> > > >
> > src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
> > > > > > > >   (revision 1560984)
> > > > > > > > > +++
> > > > > > > >
> > > > > >
> > > >
> > src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
> > > > > > > >   (working copy)
> > > > > > > > > @@ -50,6 +50,23 @@
> > > > > > > > >        + "" + CR
> > > > > > > > >        + "User-Agent: *" + CR
> > > > > > > > >        + "Disallow: /foo/bar/" + CR;   // no crawl delay for
> > > > other
> > > > > > agents
> > > > > > > > > +
> > > > > > > > > +  private static final String ROBOTS_STRING_REVERSE =
> > > > > > > > > +      "User-Agent: *" + CR
> > > > > > > > > +      + "Disallow: /foo/bar/" + CR   // no crawl delay for
> > other
> > > > > > agents
> > > > > > > > > +      + "" + CR
> > > > > > > > > +      + "User-Agent: Agent1 #foo" + CR
> > > > > > > > > +      + "Disallow: /a" + CR
> > > > > > > > > +      + "Disallow: /b/a" + CR
> > > > > > > > > +      + "#Disallow: /c" + CR
> > > > > > > > > +      + "Crawl-delay: 10" + CR  // set crawl delay for
> > Agent1
> > > > as 10
> > > > > > sec
> > > > > > > > > +      + "" + CR
> > > > > > > > > +      + "" + CR
> > > > > > > > > +      + "User-Agent: Agent2" + CR
> > > > > > > > > +      + "Disallow: /a/bloh" + CR
> > > > > > > > > +      + "Disallow: /c" + CR
> > > > > > > > > +      + "Disallow: /foo" + CR
> > > > > > > > > +      + "Crawl-delay: 20" + CR;
> > > > > > > > >
> > > > > > > > >    private static final String[] TEST_PATHS = new String[] {
> > > > > > > > >      "http://example.com/a";,
> > > > > > > > > @@ -80,6 +97,29 @@
> > > > > > > > >    /**
> > > > > > > > >    * Test that the robots rules are interpreted correctly by
> > the
> > > > > > robots
> > > > > > > > rules parser.
> > > > > > > > >    */
> > > > > > > > > +  public void testRobotsAgentReverse() {
> > > > > > > > > +    rules = parser.parseRules("testRobotsAgent",
> > > > > > > > ROBOTS_STRING_REVERSE.getBytes(), CONTENT_TYPE, SINGLE_AGENT);
> > > > > > > > > +
> > > > > > > > > +    for(int counter = 0; counter < TEST_PATHS.length;
> > > > counter++) {
> > > > > > > > > +      assertTrue("testing on agent (" + SINGLE_AGENT + "),
> > and "
> > > > > > > > > +              + "path " + TEST_PATHS[counter]
> > > > > > > > > +              + " got " +
> > rules.isAllowed(TEST_PATHS[counter]),
> > > > > > > > > +              rules.isAllowed(TEST_PATHS[counter]) ==
> > > > > > RESULTS[counter]);
> > > > > > > > > +    }
> > > > > > > > > +
> > > > > > > > > +    rules = parser.parseRules("testRobotsAgent",
> > > > > > > > ROBOTS_STRING_REVERSE.getBytes(), CONTENT_TYPE,
> > MULTIPLE_AGENTS);
> > > > > > > > > +
> > > > > > > > > +    for(int counter = 0; counter < TEST_PATHS.length;
> > > > counter++) {
> > > > > > > > > +      assertTrue("testing on agents (" + MULTIPLE_AGENTS +
> > "),
> > > > and "
> > > > > > > > > +              + "path " + TEST_PATHS[counter]
> > > > > > > > > +              + " got " +
> > rules.isAllowed(TEST_PATHS[counter]),
> > > > > > > > > +              rules.isAllowed(TEST_PATHS[counter]) ==
> > > > > > RESULTS[counter]);
> > > > > > > > > +    }
> > > > > > > > > +  }
> > > > > > > > > +
> > > > > > > > > +  /**
> > > > > > > > > +  * Test that the robots rules are interpreted correctly by
> > the
> > > > > > robots
> > > > > > > > rules parser.
> > > > > > > > > +  */
> > > > > > > > >    public void testRobotsAgent() {
> > > > > > > > >      rules = parser.parseRules("testRobotsAgent",
> > > > > > > > ROBOTS_STRING.getBytes(), CONTENT_TYPE, SINGLE_AGENT);
> > > > > > > > >
> > > > > > > > >
> > > > > > > > > -----Original message-----
> > > > > > > > > > From:Tejas Patil <[email protected]>
> > > > > > > > > > Sent: Friday 24th January 2014 15:24
> > > > > > > > > > To: [email protected]
> > > > > > > > > > Subject: Re: Order of robots file
> > > > > > > > > >
> > > > > > > > > > Hi Markus,
> > > > > > > > > > I am trying to understand the problem you described. You
> > meant
> > > > that
> > > > > > > > with
> > > > > > > > > > the original Nutch's robots parsing code, the robots file
> > below
> > > > > > allowed
> > > > > > > > > > your crawler to crawl stuff:
> > > > > > > > > >
> > > > > > > > > > User-agent: *
> > > > > > > > > > Disallow: /
> > > > > > > > > >
> > > > > > > > > > User-agent: our_crawler
> > > > > > > > > > Allow: /
> > > > > > > > > >
> > > > > > > > > > But now that started using the change from NUTCH-1031 [0],
> > (ie.
> > > > > > > > delegation
> > > > > > > > > > of robots parsing to crawler commons), it blocked your
> > > > crawler. To
> > > > > > make
> > > > > > > > > > things work, you had to change your robots file to this:
> > > > > > > > > >
> > > > > > > > > > User-agent: our_crawler
> > > > > > > > > > Allow: /
> > > > > > > > > >
> > > > > > > > > > User-agent: *
> > > > > > > > > > Disallow: /
> > > > > > > > > >
> > > > > > > > > > Did I understand the problem correctly ?
> > > > > > > > > >
> > > > > > > > > > [0] : https://issues.apache.org/jira/browse/NUTCH-1031
> > > > > > > > > >
> > > > > > > > > > Thanks,
> > > > > > > > > > Tejas
> > > > > > > > > >
> > > > > > > > > >
> > > > > > > > > > On Fri, Jan 24, 2014 at 7:29 PM, Markus Jelsma
> > > > > > > > > > <[email protected]>wrote:
> > > > > > > > > >
> > > > > > > > > > > Hi,
> > > > > > > > > > >
> > > > > > > > > > > I am attempting to merge some Nutch changes back to our
> > own.
> > > > We
> > > > > > > > aren't
> > > > > > > > > > > using Nutch' CrawlerCommons impl but the old stuff. But
> > > > because
> > > > > > of
> > > > > > > > > > > recording of response time and rudimentary SSL support i
> > > > decided
> > > > > > to
> > > > > > > > move it
> > > > > > > > > > > back to our version. Suddenly i realized a local crawl
> > does
> > > > not
> > > > > > work
> > > > > > > > > > > anymore, it seems because of the order of the robots
> > > > definitions.
> > > > > > > > > > >
> > > > > > > > > > > For example:
> > > > > > > > > > >
> > > > > > > > > > > User-agent: *
> > > > > > > > > > > Disallow: /
> > > > > > > > > > >
> > > > > > > > > > > User-agent: our_crawler
> > > > > > > > > > > Allow: /
> > > > > > > > > > >
> > > > > > > > > > > Does not allow our crawler to fetch URL's. But
> > > > > > > > > > >
> > > > > > > > > > > User-agent: our_crawler
> > > > > > > > > > > Allow: /
> > > > > > > > > > >
> > > > > > > > > > > User-agent: *
> > > > > > > > > > > Disallow: /
> > > > > > > > > > >
> > > > > > > > > > > Does! This was not the case before, anyone here aware of
> > > > this? By
> > > > > > > > design?
> > > > > > > > > > > Or is it a flaw?
> > > > > > > > > > >
> > > > > > > > > > > Thanks
> > > > > > > > > > > Markus
> > > > > > > > > > >
> > > > > > > > > >
> > > > > > > > >
> > > > > > > >
> > > > > > >
> > > > > >
> > > > >
> > > >
> > >
> >
> 

Reply via email to