You can use the host normalizer for this.
https://issues.apache.org/jira/browse/NUTCH-1319
 
-----Original message-----
> From:Arthur Yarwood <[email protected]>
> Sent: Tuesday 7th July 2015 0:02
> To: [email protected]
> Subject: Duplicate pages with and without www. prefix being indexed
> 
> I have a Nutch 1.10 and Solr 5.2 setup. Still playing around with it all 
> and quite new to me, but I've noticed for one site I have crawled, I'm 
> getting content indexed twice in Solr, once with the www. domain prefix 
> and once without. E.g.
> 
> http://www.example.com/somepage.html
> http://example.com/somepage.html
> 
> How can I avoid this duplication? At least from being indexed into Solr.
> And preferable a generic solution that will work with any other site I 
> crawl in the future, some of which may default to www.xyz.com and some 
> xyz.com. i.e. I know I could add a regex-urlfilter for this one domain, 
> but I'd like to avoid this duplication in call instances in arise.
> 
> Thanks!
> 
> 
> -- 
> Arthur Yarwood
> 
> 

Reply via email to