[jira] [Commented] (PIG-5135) HDFS bytes read stats are always 0 in Spark mode

liyunzhang_intel (JIRA) Wed, 12 Apr 2017 14:04:54 -0700

    [ 
https://issues.apache.org/jira/browse/PIG-5135?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15966615#comment-15966615
 ]


liyunzhang_intel commented on PIG-5135:
---------------------------------------

[~szita]:  I see. please create a jira board and add [~rohini] as reviewer to 
help review the modification of PigInputFormat.java.

And remove some code in org.apache.pig.test.TestPigRunner#simpleMultiQueryTest3 
as now the hdfs bytes read stats are not always 0 in spark mode
{code}
  // For mapreduce, since hdfs bytes read includes replicated tables bytes read 
is wrong
            // Since Tez does has only one load per job its values are correct
            // By pass the check for spark due to PIG-4788
            if (!Util.isMapredExecType(cluster.getExecType()) && 
!Util.isSparkExecType(cluster.getExecType())) {
                assertEquals(30, inputStats.get(0).getBytes());
                assertEquals(18, inputStats.get(1).getBytes());
            }
{code}

> HDFS bytes read stats are always 0 in Spark mode
> ------------------------------------------------
>
>                 Key: PIG-5135
>                 URL: https://issues.apache.org/jira/browse/PIG-5135
>             Project: Pig
>          Issue Type: Bug
>          Components: spark
>            Reporter: liyunzhang_intel
>            Assignee: Adam Szita
>             Fix For: spark-branch
>
>         Attachments: PIG-5135.0.patch, PIG-5135.1.patch
>
>
> I discovered this while running TestOrcStoragePushdown unit test in Spark 
> mode where the test depends on the value of this stat.



--
This message was sent by Atlassian JIRA
(v6.3.15#6346)

[jira] [Commented] (PIG-5135) HDFS bytes read stats are always 0 in Spark mode

Reply via email to