[ 
https://issues.apache.org/jira/browse/HADOOP-4565?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12669898#action_12669898
 ] 

Tom White commented on HADOOP-4565:
-----------------------------------

Dhruba,

I was thinking that the asserts may need strengthening. For example, in the 
following code (and other similar places) we should assert that the splits have 
the expected paths and locations.

{code}
// make sure that each split has different locations
for (int i = 0; i < splits.length; ++i) {
  CombineFileSplit fileSplit = (CombineFileSplit) splits[i];
  System.out.println("File split(Test1): " + fileSplit);
}
assertEquals(splits.length, 2);
{code}

Regarding HARs, I don't have a particular test in mind. We probably need a 
special input format for HARs - that would be a separate issue. 

> MultiFileInputSplit can use data locality information to create splits
> ----------------------------------------------------------------------
>
>                 Key: HADOOP-4565
>                 URL: https://issues.apache.org/jira/browse/HADOOP-4565
>             Project: Hadoop Core
>          Issue Type: Improvement
>          Components: mapred
>            Reporter: dhruba borthakur
>            Assignee: dhruba borthakur
>             Fix For: 0.20.0, 0.21.0
>
>         Attachments: CombineMultiFile.patch, CombineMultiFile2.patch, 
> CombineMultiFile3.patch, CombineMultiFile4.patch, CombineMultiFile5.patch, 
> CombineMultiFile7.patch, CombineMultiFile8.patch, CombineMultiFile9.patch
>
>
> The MultiFileInputFormat takes a set of paths and creates splits based on 
> file sizes. Each splits contains a few files an each split are roughly equal 
> in size. It would be efficient if we can extend this InputFormat to create 
> splits such each all the blocks in one split and either node-local or 
> rack-local.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.

Reply via email to