[jira] [Commented] (HIVE-15104) Hive on Spark generate more shuffle data than hive on mr

Xuefu Zhang (JIRA) Wed, 30 Aug 2017 15:48:30 -0700

    [ 
https://issues.apache.org/jira/browse/HIVE-15104?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16148174#comment-16148174
 ]


Xuefu Zhang commented on HIVE-15104:
------------------------------------

The patch looks good to me. My only concern is about the reliability of the 
runtime compilation and jar creating. I'd think it's best if we can avoid that.

I'm not 100% sure of the class loading problem we faced. If we define class 
HiveKryoRegistrator in Hive, with relocation, Spark's unrelocated kryo isn't 
able to find it?

> Hive on Spark generate more shuffle data than hive on mr
> --------------------------------------------------------
>
>                 Key: HIVE-15104
>                 URL: https://issues.apache.org/jira/browse/HIVE-15104
>             Project: Hive
>          Issue Type: Bug
>          Components: Spark
>    Affects Versions: 1.2.1
>            Reporter: wangwenli
>            Assignee: Rui Li
>         Attachments: HIVE-15104.1.patch, HIVE-15104.2.patch, 
> HIVE-15104.3.patch, HIVE-15104.4.patch, HIVE-15104.5.patch, 
> HIVE-15104.5.patch, TPC-H 100G.xlsx
>
>
> the same sql,  running on spark  and mr engine, will generate different size 
> of shuffle data.
> i think it is because of hive on mr just serialize part of HiveKey, but hive 
> on spark which using kryo will serialize full of Hivekey object.  
> what is your opionion?



--
This message was sent by Atlassian JIRA
(v6.4.14#64029)

[jira] [Commented] (HIVE-15104) Hive on Spark generate more shuffle data than hive on mr

Reply via email to