[ https://issues.apache.org/jira/browse/NUTCH-706?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12851923#action_12851923 ]
Ken Krugler commented on NUTCH-706: ----------------------------------- Two comments about this: 1. From my experiences with Nutch & Bixo, I think that URL normalization ultimately needs to be more structured - ie first break the URL into pieces, then apply rules against the pieces. Trying to craft regular expressions to handle target cases leads to big, hairy, hard-to-understand strings. 2. URL normalization is something that makes a lot of sense for crawler-commons. If somebody from the Nutch side wants to define a target API, I could look at porting existing Bixo code to crawler-commons. > Url regex normalizer > -------------------- > > Key: NUTCH-706 > URL: https://issues.apache.org/jira/browse/NUTCH-706 > Project: Nutch > Issue Type: Bug > Affects Versions: 1.0.0 > Reporter: Meghna Kukreja > Priority: Minor > > Hey, > I encountered the following problem while trying to crawl a site using > nutch-trunk. In the file regex-normalize.xml, the following regex is > used to remove session ids: > <pattern>([;_]?((?i)l|j|bv_)?((?i)sid|phpsessid|sessionid)=.*?)(\?|&|#|$)</pattern>. > This pattern also transforms a url, such as, > "&newsId=2000484784794&newsLang=en" into "&new&newsLang=en" (since it > matches 'sId' in the 'newsId'), which is incorrect and hence does not > get fetched. This expression needs to be changed to prevent this. > Thanks, > Meghna -- This message is automatically generated by JIRA. - You can reply to this email to add a comment to the issue online.