[ https://issues.apache.org/jira/browse/SOLR-799?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]
Mark Miller updated SOLR-799: ----------------------------- Attachment: SOLR-799.patch Whoops...thought I had posted this. Heres another draft I did a couple weeks ago. Fixes most of the comments brought up. There will need to be another draft - leaving the schema test file in for now as its still helpful to me. No prevent/append etc options here due to all the issues I mentioned, but I do have unposted code experimenting in the different directions if we want to try to go there anyway. > Add support for hash based exact/near duplicate document handling > ----------------------------------------------------------------- > > Key: SOLR-799 > URL: https://issues.apache.org/jira/browse/SOLR-799 > Project: Solr > Issue Type: New Feature > Components: update > Reporter: Mark Miller > Priority: Minor > Attachments: SOLR-799.patch, SOLR-799.patch > > > Hash based duplicate document detection is efficient and allows for blocking > as well as field collapsing. Lets put it into solr. > http://wiki.apache.org/solr/Deduplication -- This message is automatically generated by JIRA. - You can reply to this email to add a comment to the issue online.