In my opinion you should be able to use directly Lucene; indeed nutch
relies upon Lucene for indexing and retrieval. In your case you don't
need to crawl as files are just local and static HTML; you need to
index files and to be able to retrieve them through querying the index
so Lucene should be what you need.


-------------------------------------------------------
SF.Net email is sponsored by: Discover Easy Linux Migration Strategies
from IBM. Find simple to follow Roadmaps, straightforward articles,
informative Webcasts and more! Get everything you need to get up to
speed, fast. http://ads.osdn.com/?ad_idt77&alloc_id492&op=click
_______________________________________________
Nutch-general mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/nutch-general

Reply via email to