[EMAIL PROTECTED] wrote:
Andrzej (sorry for forgetting the 'z' earlier),

No problem - most foreigners are taken by surprise with my name... :-)

Thanks for the answer.  That makes sense and that's what I thought was
the issue - the Lucene index is just used for quick lookups.
How about de-duplication - if there are duplicate URL references in
multiple segments, which reference wins?  The most recent one?

Yes.

However, your question started me thinking... perhaps there is too much work put into indexing, deduplication, re-indexing etc. I'm trying now to write a variant of SegmentMerge which avoids most of these steps by using certain properties of Lucene indexes. This could result in radical speed improvements. Stay tuned.

--
Best regards,
Andrzej Bialecki

-------------------------------------------------
Software Architect, System Integration Specialist
CEN/ISSS EC Workshop, ECIMF project chair
EU FP6 E-Commerce Expert/Evaluator
-------------------------------------------------
FreeBSD developer (http://www.freebsd.org)



-------------------------------------------------------
This SF.net email is sponsored by: IT Product Guide on ITManagersJournal
Use IT products in your business? Tell us what you think of them. Give us
Your Opinions, Get Free ThinkGeek Gift Certificates! Click to find out more
http://productguide.itmanagersjournal.com/guidepromo.tmpl
_______________________________________________
Nutch-developers mailing list
[EMAIL PROTECTED]
https://lists.sourceforge.net/lists/listinfo/nutch-developers

Reply via email to