[jira] [Commented] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records

Wilfred Spiegelenburg (JIRA) Fri, 13 May 2016 21:44:24 -0700

    [ 
https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15283419#comment-15283419
 ]


Wilfred Spiegelenburg commented on MAPREDUCE-6558:
--------------------------------------------------

I somehow could not leave a comment yesterday. I made the .3 patch to fix some 
comments in the test code and decrease the size of the test file even further.
Thank you for the review and the commit [~jlowe]


> multibyte delimiters with compressed input files generate duplicate records
> ---------------------------------------------------------------------------
>
>                 Key: MAPREDUCE-6558
>                 URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558
>             Project: Hadoop Map/Reduce
>          Issue Type: Bug
>          Components: mrv1, mrv2
>    Affects Versions: 2.7.2
>            Reporter: Wilfred Spiegelenburg
>            Assignee: Wilfred Spiegelenburg
>             Fix For: 2.8.0, 2.7.3, 2.6.5
>
>         Attachments: MAPREDUCE-6558.1.patch, MAPREDUCE-6558.2.patch, 
> MAPREDUCE-6558.3.patch
>
>
> This is the follow up for MAPREDUCE-6549. Compressed files cause record 
> duplications as shown in different junit tests. The number of duplicated 
> records changes with the splitsize:
> Unexpected number of records in split (splitsize = 10)
> Expected: 41051
> Actual: 45062
> Unexpected number of records in split (splitsize = 100000)
> Expected: 41051
> Actual: 41052
> Test passes with splitsize = 147445 which is the compressed file length.The 
> file is a bzip2 file with 100k blocks and a total of 11 blocks



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

[jira] [Commented] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records

Reply via email to