Bikas Saha created MAPREDUCE-4892:
-------------------------------------

             Summary: CombineFileInputFormat node input split can be skewed on 
small clusters
                 Key: MAPREDUCE-4892
                 URL: https://issues.apache.org/jira/browse/MAPREDUCE-4892
             Project: Hadoop Map/Reduce
          Issue Type: Bug
            Reporter: Bikas Saha
            Assignee: Bikas Saha
             Fix For: 3.0.0


The CombineFileInputFormat split generation logic tries to group blocks by node 
in order to create splits. It iterates through the nodes and creates splits on 
them until there aren't enough blocks left on a node that can be grouped into a 
valid split. If the first few nodes have a lot of blocks on them then they can 
end up getting a disproportionately large share of the total number of splits 
created. This can result in poor locality of maps. This problem is likely to 
happen on small clusters where its easier to create a skew in the distribution 
of blocks on nodes.

--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira

Reply via email to