Hi, sorry for being unclear. You understand correctly, i had to change the
robots.txt order and put our crawler ABOVE User-Agent: *.
I have tried a unit test for lib-http to demonstrate the problem but i does not
fail. I think i did something wrong in merging the code base. I'll look further.
markus@midas:~/projects/apache/nutch/trunk$ svn diff
src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
Index:
src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
===================================================================
---
src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
(revision 1560984)
+++
src/plugin/lib-http/src/test/org/apache/nutch/protocol/http/api/TestRobotRulesParser.java
(working copy)
@@ -50,6 +50,23 @@
+ "" + CR
+ "User-Agent: *" + CR
+ "Disallow: /foo/bar/" + CR; // no crawl delay for other agents
+
+ private static final String ROBOTS_STRING_REVERSE =
+ "User-Agent: *" + CR
+ + "Disallow: /foo/bar/" + CR // no crawl delay for other agents
+ + "" + CR
+ + "User-Agent: Agent1 #foo" + CR
+ + "Disallow: /a" + CR
+ + "Disallow: /b/a" + CR
+ + "#Disallow: /c" + CR
+ + "Crawl-delay: 10" + CR // set crawl delay for Agent1 as 10 sec
+ + "" + CR
+ + "" + CR
+ + "User-Agent: Agent2" + CR
+ + "Disallow: /a/bloh" + CR
+ + "Disallow: /c" + CR
+ + "Disallow: /foo" + CR
+ + "Crawl-delay: 20" + CR;
private static final String[] TEST_PATHS = new String[] {
"http://example.com/a",
@@ -80,6 +97,29 @@
/**
* Test that the robots rules are interpreted correctly by the robots rules
parser.
*/
+ public void testRobotsAgentReverse() {
+ rules = parser.parseRules("testRobotsAgent",
ROBOTS_STRING_REVERSE.getBytes(), CONTENT_TYPE, SINGLE_AGENT);
+
+ for(int counter = 0; counter < TEST_PATHS.length; counter++) {
+ assertTrue("testing on agent (" + SINGLE_AGENT + "), and "
+ + "path " + TEST_PATHS[counter]
+ + " got " + rules.isAllowed(TEST_PATHS[counter]),
+ rules.isAllowed(TEST_PATHS[counter]) == RESULTS[counter]);
+ }
+
+ rules = parser.parseRules("testRobotsAgent",
ROBOTS_STRING_REVERSE.getBytes(), CONTENT_TYPE, MULTIPLE_AGENTS);
+
+ for(int counter = 0; counter < TEST_PATHS.length; counter++) {
+ assertTrue("testing on agents (" + MULTIPLE_AGENTS + "), and "
+ + "path " + TEST_PATHS[counter]
+ + " got " + rules.isAllowed(TEST_PATHS[counter]),
+ rules.isAllowed(TEST_PATHS[counter]) == RESULTS[counter]);
+ }
+ }
+
+ /**
+ * Test that the robots rules are interpreted correctly by the robots rules
parser.
+ */
public void testRobotsAgent() {
rules = parser.parseRules("testRobotsAgent", ROBOTS_STRING.getBytes(),
CONTENT_TYPE, SINGLE_AGENT);
-----Original message-----
> From:Tejas Patil <[email protected]>
> Sent: Friday 24th January 2014 15:24
> To: [email protected]
> Subject: Re: Order of robots file
>
> Hi Markus,
> I am trying to understand the problem you described. You meant that with
> the original Nutch's robots parsing code, the robots file below allowed
> your crawler to crawl stuff:
>
> User-agent: *
> Disallow: /
>
> User-agent: our_crawler
> Allow: /
>
> But now that started using the change from NUTCH-1031 [0], (ie. delegation
> of robots parsing to crawler commons), it blocked your crawler. To make
> things work, you had to change your robots file to this:
>
> User-agent: our_crawler
> Allow: /
>
> User-agent: *
> Disallow: /
>
> Did I understand the problem correctly ?
>
> [0] : https://issues.apache.org/jira/browse/NUTCH-1031
>
> Thanks,
> Tejas
>
>
> On Fri, Jan 24, 2014 at 7:29 PM, Markus Jelsma
> <[email protected]>wrote:
>
> > Hi,
> >
> > I am attempting to merge some Nutch changes back to our own. We aren't
> > using Nutch' CrawlerCommons impl but the old stuff. But because of
> > recording of response time and rudimentary SSL support i decided to move it
> > back to our version. Suddenly i realized a local crawl does not work
> > anymore, it seems because of the order of the robots definitions.
> >
> > For example:
> >
> > User-agent: *
> > Disallow: /
> >
> > User-agent: our_crawler
> > Allow: /
> >
> > Does not allow our crawler to fetch URL's. But
> >
> > User-agent: our_crawler
> > Allow: /
> >
> > User-agent: *
> > Disallow: /
> >
> > Does! This was not the case before, anyone here aware of this? By design?
> > Or is it a flaw?
> >
> > Thanks
> > Markus
> >
>