Hi nutcher's, ;-) http://www.media-style.com/gfx/assets/nutch_character_insg1.gif (some action pictures )
I would love to discuss about url filtering / normalizing with you. If I do miss understand or oversee something - please correct me. We actually have a regular expression injection url filter. This filter is use as well in the db update tool.
I would love to re-factor the code to a Extension point and a plugin.
Wouldn't it be sense-full to use the filter in the fetching process? Does the fetcher not follow links directly?
Wouldn't it be useful to have not fetched links in the db as well but just not fetch it?
Isn't a page that have more links to other important pages may be more interesting then just normal pages?
So the google principle reverse with may be a lower weight?
If I try to evaluation a science paper i take at first a look at the quotations. Does Nutch already do that, I'm don't know?
For url filtering I have following things in mind.
Filtering based on:
1.) dns whois queries. (http://jakarta.apache.org/commons/net/)
2.) Ip2 geo position database queries (for example http://www.ip2location.com/)
3.) link extraction based on content, for example: keyword - domain thesaurus matching, text classification or ontology.
Any 2 cents?
Regards, Stefan Groschupf
--------------------------------------------------------------- enterprise information technology consultanting open technology: http://www.media-style.com open source: http://www.weta-group.net open discussion: http://www.text-mining.org
------------------------------------------------------- This SF.Net email is sponsored by the new InstallShield X.
From Windows to Linux, servers to mobile, InstallShield X is the one
installation-authoring solution that does it all. Learn more and evaluate today! http://www.installshield.com/Dev2Dev/0504 _______________________________________________ Nutch-developers mailing list [EMAIL PROTECTED] https://lists.sourceforge.net/lists/listinfo/nutch-developers
