On 2/9/06, Stefan Groschupf <[EMAIL PROTECTED]> wrote:
> Hi Folks,
>
> I hope and it looks like we are close to get meta data support for
> crawlDatum (CrawlDB) into the sources soon.
> At this point we can store and read but not 'process' (means creation
> or inheritance etc. [some one knows a better naming for this
> process]) meta data. So may plan is to add some plugin extension
> points that allows to plug - in meta data processor plugins.
>
> There are so many use cases for meta data usage, that I would love to
> ask you for any suggestion where you think we should add support for
> meta data processing.
> As soon a plugin can manipulate a crawlDatum it also can process
> somehow meta data by read and write meta data.  e.g.
> crawlDatum.getMetaData().get(KEY) or getMetaData().put(KEY, myValue).
> So in such cases we don't need to change anything.
>
> However a use case could be to hand over some meta data from a mother
> page to a child page. e.g. crawler policy or a category.
> So the limit is just the creativity ... and the performance. That is
> way we should try to discuss what are the hot spots for meta data
> processing today.
>
>
> So let me start with my personal hotspots.
>
> I actually already have a injector  that supports meta data
> injection. This injector can parse a url file where we have first the
> url followed by key:value TAB key:value meta data tuples.
> Also I have a index filter plugin that reads store meta data as
> key>fieldName, value->keyword into the index.
>
> Beside that I would love to see the possibility to hand over meta
> data (done by a plugin) from a mother page to a page that created
> from the links found on the 'mother' page.

Hi Stefan.
Yes, I am facing the problem on creating meta-data from distribted pages.
Say, page A<url=ps.html?id=123> contains "price" meta data, while
page B<pd.html?id=123> contains "description" meta data. Is it possible
to extract meta data from "page A join page B" if the plugin extends HtmlParser?

And now my trying is extracting the metadata when re-visit
webdb(offline). I mean we should segment the meta data. One kind can
be only extracted online while another can
be done offline. Comments?

/Jack

> So please add your hotspots and let's discuss where and how we can
> handling meta data to not slow down things in nutch.
>
> Thanks for thoughts.
> Stefan
>


--
Keep Discovering ... ...
http://www.jroller.com/page/jmars

Reply via email to