[
https://issues.apache.org/jira/browse/NUTCH-1149?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Markus Jelsma closed NUTCH-1149.
--------------------------------
Resolution: Won't Fix
Will upload proper patch for NUTCH-1325 soon which already contains numeric
aggregations for CrawlDB metadata.
> DomainStats should process numeric CrawlDB metadata
> ---------------------------------------------------
>
> Key: NUTCH-1149
> URL: https://issues.apache.org/jira/browse/NUTCH-1149
> Project: Nutch
> Issue Type: Improvement
> Reporter: Markus Jelsma
> Assignee: Markus Jelsma
> Priority: Trivial
>
> Right now the DomainStats program only outputs the sum of fetched records per
> domain or host. It should also be able to output processed numerics of meta
> data in order to get the average size (content length) for a given domain or
> host. This is also useful for generating a metric for adult material (by
> domain or host) when using a plugin that stores a propability factor of adult
> material per URL in the Crawl DB.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)