[ https://issues.apache.org/jira/browse/NUTCH-1314?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15137467#comment-15137467 ]
Hudson commented on NUTCH-1314: ------------------------------- SUCCESS: Integrated in Nutch-nutchgora #1549 (See [https://builds.apache.org/job/Nutch-nutchgora/1549/]) NUTCH-1314 Impose a limit on the length of outlink target urls (lewismc: [http://svn.apache.org/viewvc/nutch/branches/2.x/?view=rev&rev=1729220]) * 2.x/conf/nutch-default.xml NUTCH-1314 Impose a limit on the length of outlink target urls (lewismc: [http://svn.apache.org/viewvc/nutch/branches/2.x/?view=rev&rev=1729219]) * 2.x/src/test/org/apache/nutch/parse/TestParseUtil.java NUTCH-1314 Impose a limit on the length of outlink target urls (lewismc: [http://svn.apache.org/viewvc/nutch/branches/2.x/?view=rev&rev=1729218]) * 2.x/CHANGES.txt * 2.x/conf/nutch-default.xml * 2.x/src/java/org/apache/nutch/parse/ParseUtil.java > Impose a limit on the length of outlink target urls > --------------------------------------------------- > > Key: NUTCH-1314 > URL: https://issues.apache.org/jira/browse/NUTCH-1314 > Project: Nutch > Issue Type: Improvement > Reporter: Ferdy Galema > Assignee: Lewis John McGibbney > Fix For: 2.4, 1.12 > > Attachments: NUTCH-1314-trunk.patch, NUTCH-1314-v2.patch, > NUTCH-1314-v3.patch, NUTCH-1314-v4.patch, NUTCH-1314.patch > > > In the past we have encountered situations where crawling specific broken > sites resulted in ridiciously long urls that caused the stalling of tasks. > The regex plugins (normalizing/filtering) processed single urls for hours, if > not indefinitely hanging. > My suggestion is to limit the outlink url target length as soon possible. It > is a configurable limit, the default is 3000. This should be reasonably long > enough for most uses. But sufficienly strict enough to make sure regex > plugins do not choke on urls that are too long. Please see attached patch for > the Nutchgora implementation. > I'd like to hear what you think about this. -- This message was sent by Atlassian JIRA (v6.3.4#6332)