[
https://issues.apache.org/jira/browse/HADOOP-3514?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Arun C Murthy updated HADOOP-3514:
----------------------------------
Status: Open (was: Patch Available)
Some comments:
1. Checksum{Input|Output}Stream shouldn't be public - in fact I'd urge they be
renamed to IFile{Input|Output}Stream, to emphasize their 'internal' nature.
2. I'm a little bit tetchy about ChecksumInputStream's behaviour wrt
'remembering' that the last 4 bytes are the checksum and the complications it
brings along. I'd propose a simpler alternative where we store the checksum in
the index-file along with offset/compressed-length/decompressed-length. Then we
could send over the checksum via the http headers, pass the received checksum
along to the constructor of ChecksumInptuStream which would then validate the
running-checksum against the stored checksum in the 'close()' method - similar
to how it is produced in the ChecksumOutputStream's close method.
Thoughts?
> Reduce seeks during shuffle, by inline crcs
> -------------------------------------------
>
> Key: HADOOP-3514
> URL: https://issues.apache.org/jira/browse/HADOOP-3514
> Project: Hadoop Core
> Issue Type: Improvement
> Components: mapred
> Affects Versions: 0.18.0
> Reporter: Devaraj Das
> Assignee: Jothi Padmanabhan
> Fix For: 0.19.0
>
> Attachments: hadoop-3514-v1.patch, hadoop-3514-v2.patch,
> hadoop-3514-v3.patch, hadoop-3514-v4.patch, hadoop-3514-v5.patch,
> hadoop-3514-v6.patch, hadoop-3514-v7.patch, hadoop-3514-v8.patch,
> hadoop-3514.patch
>
>
> The number of seeks can be reduced by half in the iFile if we move the crc
> into the iFile rather than having a separate file.
--
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.