In my opinion you should be able to use directly Lucene; indeed nutch relies upon Lucene for indexing and retrieval. In your case you don't need to crawl as files are just local and static HTML; you need to index files and to be able to retrieve them through querying the index so Lucene should be what you need.
------------------------------------------------------- SF.Net email is sponsored by: Discover Easy Linux Migration Strategies from IBM. Find simple to follow Roadmaps, straightforward articles, informative Webcasts and more! Get everything you need to get up to speed, fast. http://ads.osdn.com/?ad_idt77&alloc_id492&op=click _______________________________________________ Nutch-general mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/nutch-general
