On 2/9/06, Stefan Groschupf <[EMAIL PROTECTED]> wrote: > Hi Folks, > > I hope and it looks like we are close to get meta data support for > crawlDatum (CrawlDB) into the sources soon. > At this point we can store and read but not 'process' (means creation > or inheritance etc. [some one knows a better naming for this > process]) meta data. So may plan is to add some plugin extension > points that allows to plug - in meta data processor plugins. > > There are so many use cases for meta data usage, that I would love to > ask you for any suggestion where you think we should add support for > meta data processing. > As soon a plugin can manipulate a crawlDatum it also can process > somehow meta data by read and write meta data. e.g. > crawlDatum.getMetaData().get(KEY) or getMetaData().put(KEY, myValue). > So in such cases we don't need to change anything. > > However a use case could be to hand over some meta data from a mother > page to a child page. e.g. crawler policy or a category. > So the limit is just the creativity ... and the performance. That is > way we should try to discuss what are the hot spots for meta data > processing today. > > > So let me start with my personal hotspots. > > I actually already have a injector that supports meta data > injection. This injector can parse a url file where we have first the > url followed by key:value TAB key:value meta data tuples. > Also I have a index filter plugin that reads store meta data as > key>fieldName, value->keyword into the index. > > Beside that I would love to see the possibility to hand over meta > data (done by a plugin) from a mother page to a page that created > from the links found on the 'mother' page.
Hi Stefan. Yes, I am facing the problem on creating meta-data from distribted pages. Say, page A<url=ps.html?id=123> contains "price" meta data, while page B<pd.html?id=123> contains "description" meta data. Is it possible to extract meta data from "page A join page B" if the plugin extends HtmlParser? And now my trying is extracting the metadata when re-visit webdb(offline). I mean we should segment the meta data. One kind can be only extracted online while another can be done offline. Comments? /Jack > So please add your hotspots and let's discuss where and how we can > handling meta data to not slow down things in nutch. > > Thanks for thoughts. > Stefan > -- Keep Discovering ... ... http://www.jroller.com/page/jmars
