[ 
https://issues.apache.org/jira/browse/HADOOP-3514?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Arun C Murthy updated HADOOP-3514:
----------------------------------

    Status: Open  (was: Patch Available)

Some comments:

1. Checksum{Input|Output}Stream shouldn't be public - in fact I'd urge they be 
renamed to IFile{Input|Output}Stream, to emphasize their 'internal' nature.
2. I'm a little bit tetchy about ChecksumInputStream's behaviour wrt 
'remembering' that the last 4 bytes are the checksum and the complications it 
brings along. I'd propose a simpler alternative where we store the checksum in 
the index-file along with offset/compressed-length/decompressed-length. Then we 
could send over the checksum via the http headers, pass the received checksum 
along to the constructor of ChecksumInptuStream which would then validate the 
running-checksum against the stored checksum in the 'close()' method - similar 
to how it is produced in the ChecksumOutputStream's close method. 

Thoughts?

> Reduce seeks during shuffle, by inline crcs
> -------------------------------------------
>
>                 Key: HADOOP-3514
>                 URL: https://issues.apache.org/jira/browse/HADOOP-3514
>             Project: Hadoop Core
>          Issue Type: Improvement
>          Components: mapred
>    Affects Versions: 0.18.0
>            Reporter: Devaraj Das
>            Assignee: Jothi Padmanabhan
>             Fix For: 0.19.0
>
>         Attachments: hadoop-3514-v1.patch, hadoop-3514-v2.patch, 
> hadoop-3514-v3.patch, hadoop-3514-v4.patch, hadoop-3514-v5.patch, 
> hadoop-3514-v6.patch, hadoop-3514-v7.patch, hadoop-3514-v8.patch, 
> hadoop-3514.patch
>
>
> The number of seeks can be reduced by half in the iFile if we move the crc 
> into the iFile rather than having a separate file.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.

Reply via email to