[
https://issues.apache.org/jira/browse/NUTCH-1314?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13694784#comment-13694784
]
Canan Girgin commented on NUTCH-1314:
-------------------------------------
I tried to test NUTCH-1314-v2.patch. But it removes links size<3000.In my
opinion, "if (target.length() > maxTargetLength)" rows are not correct in patch
file. It must be like "if (target.length() < maxTargetLength) ".
NUTCH-1314-v2.patch file , there is a new parameter used
("parser.html.outlinks.max_target_length"). I think it must be defined in
nutch-default.xml file.
I attached a new patch file. In the ParseUtil class, target url length
controlled before normalizer and filters. Is it correct?
> Impose a limit on the length of outlink target urls
> ---------------------------------------------------
>
> Key: NUTCH-1314
> URL: https://issues.apache.org/jira/browse/NUTCH-1314
> Project: Nutch
> Issue Type: Improvement
> Reporter: Ferdy Galema
> Fix For: 2.3
>
> Attachments: NUTCH-1314.patch, NUTCH-1314-trunk.patch,
> NUTCH-1314-v2.patch, NUTCH-1314-v3.patch
>
>
> In the past we have encountered situations where crawling specific broken
> sites resulted in ridiciously long urls that caused the stalling of tasks.
> The regex plugins (normalizing/filtering) processed single urls for hours, if
> not indefinitely hanging.
> My suggestion is to limit the outlink url target length as soon possible. It
> is a configurable limit, the default is 3000. This should be reasonably long
> enough for most uses. But sufficienly strict enough to make sure regex
> plugins do not choke on urls that are too long. Please see attached patch for
> the Nutchgora implementation.
> I'd like to hear what you think about this.
--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira