[ https://issues.apache.org/jira/browse/MAPREDUCE-6166?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]
Vinod Kumar Vavilapalli updated MAPREDUCE-6166: ----------------------------------------------- Target Version/s: 2.7.0, 3.0.0 (was: 3.0.0, 2.7.0) Fix Version/s: 2.6.1 Pulled this into 2.6.1. Ran compilation and TestFetcher before the push. Patch applied cleanly. > Reducers do not validate checksum of map outputs when fetching directly to > disk > ------------------------------------------------------------------------------- > > Key: MAPREDUCE-6166 > URL: https://issues.apache.org/jira/browse/MAPREDUCE-6166 > Project: Hadoop Map/Reduce > Issue Type: Bug > Components: mrv2 > Affects Versions: 2.6.0 > Reporter: Eric Payne > Assignee: Eric Payne > Labels: 2.6.1-candidate > Fix For: 2.7.0, 2.6.1 > > Attachments: MAPREDUCE-6166.v1.201411221941.txt, > MAPREDUCE-6166.v2.201411251627.txt, MAPREDUCE-6166.v3.txt, > MAPREDUCE-6166.v4.txt, MAPREDUCE-6166.v5.txt > > > In very large map/reduce jobs (50000 maps, 2500 reducers), the intermediate > map partition output gets corrupted on disk on the map side. If this > corrupted map output is too large to shuffle in memory, the reducer streams > it to disk without validating the checksum. In jobs this large, it could take > hours before the reducer finally tries to read the corrupted file and fails. > Since retries of the failed reduce attempt will also take hours, this delay > in discovering the failure is multiplied greatly. -- This message was sent by Atlassian JIRA (v6.3.4#6332)