[ https://issues.apache.org/jira/browse/HADOOP-4565?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12857586#action_12857586 ]
Gaurav Jain commented on HADOOP-4565: ------------------------------------- Did you happen to hit the following issue after this jiras changes? https://issues.apache.org/jira/browse/HDFS-347 Looks like I am hitting this issue. I am running some more benchmarks to trace down the details. > MultiFileInputSplit can use data locality information to create splits > ---------------------------------------------------------------------- > > Key: HADOOP-4565 > URL: https://issues.apache.org/jira/browse/HADOOP-4565 > Project: Hadoop Common > Issue Type: Improvement > Reporter: dhruba borthakur > Assignee: dhruba borthakur > Fix For: 0.20.0 > > Attachments: CombineMultiFile.patch, CombineMultiFile2.patch, > CombineMultiFile3.patch, CombineMultiFile4.patch, CombineMultiFile5.patch, > CombineMultiFile7.patch, CombineMultiFile8.patch, CombineMultiFile9.patch, > TestCombine.txt > > > The MultiFileInputFormat takes a set of paths and creates splits based on > file sizes. Each splits contains a few files an each split are roughly equal > in size. It would be efficient if we can extend this InputFormat to create > splits such each all the blocks in one split and either node-local or > rack-local. -- This message is automatically generated by JIRA. - If you think it was sent incorrectly contact one of the administrators: https://issues.apache.org/jira/secure/Administrators.jspa - For more information on JIRA, see: http://www.atlassian.com/software/jira