[jira] [Commented] (NUTCH-2155) Create a "crawl completeness" utility

ASF GitHub Bot (JIRA) Wed, 28 Oct 2015 15:04:10 -0700

    [ 
https://issues.apache.org/jira/browse/NUTCH-2155?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14979325#comment-14979325
 ]


ASF GitHub Bot commented on NUTCH-2155:
---------------------------------------

Github user MJJoyce commented on a diff in the pull request:

    https://github.com/apache/nutch/pull/83#discussion_r43324656
  
    --- Diff: src/java/org/apache/nutch/util/CrawlCompletionStats.java ---
    @@ -0,0 +1,189 @@
    +/**
    + * Licensed to the Apache Software Foundation (ASF) under one or more
    + * contributor license agreements.  See the NOTICE file distributed with
    + * this work for additional information regarding copyright ownership.
    + * The ASF licenses this file to You under the Apache License, Version 2.0
    + * (the "License"); you may not use this file except in compliance with
    + * the License.  You may obtain a copy of the License at
    + *
    + *     http://www.apache.org/licenses/LICENSE-2.0
    + *
    + * Unless required by applicable law or agreed to in writing, software
    + * distributed under the License is distributed on an "AS IS" BASIS,
    + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
    + * See the License for the specific language governing permissions and
    + * limitations under the License.
    + */
    +
    +package org.apache.nutch.util;
    +
    +import java.io.IOException;
    +import java.net.URL;
    +import java.text.SimpleDateFormat;
    +import org.slf4j.Logger;
    +import org.slf4j.LoggerFactory;
    +import org.apache.hadoop.conf.Configuration;
    +import org.apache.hadoop.conf.Configured;
    +import org.apache.hadoop.fs.Path;
    +import org.apache.hadoop.io.LongWritable;
    +import org.apache.hadoop.io.Text;
    +import org.apache.hadoop.mapreduce.Job;
    +import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
    +import org.apache.hadoop.mapreduce.lib.input.SequenceFileInputFormat;
    +import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;
    +import org.apache.hadoop.mapreduce.lib.output.TextOutputFormat;
    +import org.apache.hadoop.mapreduce.Mapper;
    +import org.apache.hadoop.mapreduce.Reducer;
    +import org.apache.hadoop.util.Tool;
    +import org.apache.hadoop.util.ToolRunner;
    +import org.apache.nutch.crawl.CrawlDatum;
    +import org.apache.nutch.util.NutchConfiguration;
    +import org.apache.nutch.util.TimingUtil;
    +import org.apache.nutch.util.URLUtil;
    +
    +/**
    + * Extracts some simple crawl completion stats from the crawldb
    + *
    + * Stats will be sorted by host/domain and will be of the form:
    + * 1       www.spitzer.caltech.edu FETCHED
    + * 50      www.spitzer.caltech.edu UNFETCHED
    + *
    + */
    +public class CrawlCompletionStats extends Configured implements Tool {
    +
    +  private static final Logger LOG = LoggerFactory
    +      .getLogger(CrawlCompletionStats.class);
    +
    +  private static final int MODE_HOST = 1;
    +  private static final int MODE_DOMAIN = 2;
    +
    +  private int mode = 0;
    +
    +  public int run(String[] args) throws Exception {
    +    if (args.length < 2) {
    --- End diff --
    
    +1 I agree completely @lewismc. I got a bit lazy and stole some from 
domainstats (which is also in need of some commons-cli love as well). I'll try 
to throw a patch together an address some of these issues when I get some free 
time.


> Create a "crawl completeness" utility
> -------------------------------------
>
>                 Key: NUTCH-2155
>                 URL: https://issues.apache.org/jira/browse/NUTCH-2155
>             Project: Nutch
>          Issue Type: Improvement
>          Components: util
>    Affects Versions: 1.10
>            Reporter: Michael Joyce
>             Fix For: 1.12
>
>
> I've found it useful to have a tool for dumping some "completeness" 
> information from a crawl similar to how domainstats does but including 
> fetched and unfetched counts per domain/host. This is especially nice when 
> doing vertical crawls over a few domains or just to see how much of a 
> host/domain you've covered with your crawl so far.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

[jira] [Commented] (NUTCH-2155) Create a "crawl completeness" utility

Reply via email to