Hi, sorry for being unclear. You understand correctly, i had to change the 
robots.txt order and put our crawler ABOVE User-Agent: *.

I have tried a unit test for lib-http to demonstrate the problem but i does not 
fail. I think i did something wrong in merging the code base. I'll look further.

markus@midas:~/projects/apache/nutch/trunk$ svn diff 
src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
Index: 
src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
===================================================================
--- 
src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
   (revision 1560984)
+++ 
src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
   (working copy)
@@ -50,6 +50,23 @@
       + "" + CR 
       + "User-Agent: *" + CR 
       + "Disallow: /foo/bar/" + CR;   // no crawl delay for other agents
+      
+  private static final String ROBOTS_STRING_REVERSE = 
+      "User-Agent: *" + CR 
+      + "Disallow: /foo/bar/" + CR   // no crawl delay for other agents
+      + "" + CR 
+      + "User-Agent: Agent1 #foo" + CR 
+      + "Disallow: /a" + CR 
+      + "Disallow: /b/a" + CR 
+      + "#Disallow: /c" + CR 
+      + "Crawl-delay: 10" + CR  // set crawl delay for Agent1 as 10 sec
+      + "" + CR 
+      + "" + CR 
+      + "User-Agent: Agent2" + CR 
+      + "Disallow: /a/bloh" + CR 
+      + "Disallow: /c" + CR
+      + "Disallow: /foo" + CR
+      + "Crawl-delay: 20" + CR;
   
   private static final String[] TEST_PATHS = new String[] {
     "http://example.com/a";,
@@ -80,6 +97,29 @@
   /**
   * Test that the robots rules are interpreted correctly by the robots rules 
parser. 
   */
+  public void testRobotsAgentReverse() {
+    rules = parser.parseRules("testRobotsAgent", 
ROBOTS_STRING_REVERSE.getBytes(), CONTENT_TYPE, SINGLE_AGENT);
+
+    for(int counter = 0; counter < TEST_PATHS.length; counter++) {
+      assertTrue("testing on agent (" + SINGLE_AGENT + "), and " 
+              + "path " + TEST_PATHS[counter] 
+              + " got " + rules.isAllowed(TEST_PATHS[counter]),
+              rules.isAllowed(TEST_PATHS[counter]) == RESULTS[counter]);
+    }
+
+    rules = parser.parseRules("testRobotsAgent", 
ROBOTS_STRING_REVERSE.getBytes(), CONTENT_TYPE, MULTIPLE_AGENTS);
+
+    for(int counter = 0; counter < TEST_PATHS.length; counter++) {
+      assertTrue("testing on agents (" + MULTIPLE_AGENTS + "), and " 
+              + "path " + TEST_PATHS[counter] 
+              + " got " + rules.isAllowed(TEST_PATHS[counter]),
+              rules.isAllowed(TEST_PATHS[counter]) == RESULTS[counter]);
+    }
+  }
+  
+  /**
+  * Test that the robots rules are interpreted correctly by the robots rules 
parser. 
+  */
   public void testRobotsAgent() {
     rules = parser.parseRules("testRobotsAgent", ROBOTS_STRING.getBytes(), 
CONTENT_TYPE, SINGLE_AGENT);

 
-----Original message-----
> From:Tejas Patil <[email protected]>
> Sent: Friday 24th January 2014 15:24
> To: [email protected]
> Subject: Re: Order of robots file
> 
> Hi Markus,
> I am trying to understand the problem you described. You meant that with
> the original Nutch's robots parsing code, the robots file below allowed
> your crawler to crawl stuff:
> 
> User-agent: *
> Disallow: /
> 
> User-agent: our_crawler
> Allow: /
> 
> But now that started using the change from NUTCH-1031 [0], (ie. delegation
> of robots parsing to crawler commons), it blocked your crawler. To make
> things work, you had to change your robots file to this:
> 
> User-agent: our_crawler
> Allow: /
> 
> User-agent: *
> Disallow: /
> 
> Did I understand the problem correctly ?
> 
> [0] : https://issues.apache.org/jira/browse/NUTCH-1031
> 
> Thanks,
> Tejas
> 
> 
> On Fri, Jan 24, 2014 at 7:29 PM, Markus Jelsma
> <[email protected]>wrote:
> 
> > Hi,
> >
> > I am attempting to merge some Nutch changes back to our own. We aren't
> > using Nutch' CrawlerCommons impl but the old stuff. But because of
> > recording of response time and rudimentary SSL support i decided to move it
> > back to our version. Suddenly i realized a local crawl does not work
> > anymore, it seems because of the order of the robots definitions.
> >
> > For example:
> >
> > User-agent: *
> > Disallow: /
> >
> > User-agent: our_crawler
> > Allow: /
> >
> > Does not allow our crawler to fetch URL's. But
> >
> > User-agent: our_crawler
> > Allow: /
> >
> > User-agent: *
> > Disallow: /
> >
> > Does! This was not the case before, anyone here aware of this? By design?
> > Or is it a flaw?
> >
> > Thanks
> > Markus
> >
> 

Reply via email to