If working on near-duplicate page duplication, here are a couple more papers that might be of interest to you.
http://www.ir.iit.edu/~abdur/publications/470-kolcz.pdf http://www.ir.iit.edu/publications/downloads/p171-chowdhury.pdf On 6/19/06, Andrzej Bialecki <[EMAIL PROTECTED]> wrote:
Eugen Kochuev wrote: > Hello nutch-user, > > I'm going to implement algorithm of filtering out near duplicates of > the web pages. The classic solution here is shingles algorithm > http://gatekeeper.research.compaq.com/pub/DEC/SRC/technical-notes/SRC-1997-015-html/ > That would be a nice addition, I agree. According to the section "The clustering algorithm" in the above-mentioned paper, it seems that you will need at least two map-reduce jobs. The first job: * map: - input: <url, parseText>. - output: <shingle, url>* * reduce: - input: <shingle, url*> - output: <shingle, <url-n, url-m>> for every pair of unique input urls. If there's only 1 input url, skip this record altogether. The second job: * map: - input: <shingle, <url-n, url-m>> - output: <<url-n, url-m>, 1> * reduce: - input: <<url-n, url-m>, 1>* - output: <count, <url-n, url-m>> Since the output of the last reduce job is sorted by count (you should use your own Comparator to reverse-sort it, so that the numbers decrease), you automatically get a list <url-n, url-m> pairs that share the largest number of shingles. Finally, you simply read this list until the counts fall below a threshold, and then you consider all pages above threshold to be duplicates, and keep only one of each pair. -- Best regards, Andrzej Bialecki <>< ___. ___ ___ ___ _ _ __________________________________ [__ || __|__/|__||\/| Information Retrieval, Semantic Web ___|||__|| \| || | Embedded Unix, System Integration http://www.sigram.com Contact: info at sigram dot com
_______________________________________________ Nutch-general mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/nutch-general
