Hi,

during a long fetch-run I experienced session-IDs in URLs, which was a
bit problematic. So I figured out how to write and test proper
regex-normalizer-rules (see NUTCH-279).

Now I wonder if on the next fetch-round URLs will get properly
normalized of if they are now un-normalized in the crawldb and from
there are fetched during generate without realizing the "duplicate"
(after normalization) URLs.

Also, is there a way to "clean" the page-index before actually indexing?
Our would this automatically be taken care of (does the normalizere run
again?) when performing the actual invertlinks/index/dedup?


Regards,
 Stefan


-------------------------------------------------------
Using Tomcat but need to do more? Need to support web services, security?
Get stuff done quickly with pre-integrated technology to make your job easier
Download IBM WebSphere Application Server v.1.0.1 based on Apache Geronimo
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=120709&bid=263057&dat=121642
_______________________________________________
Nutch-general mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/nutch-general

Reply via email to