[
https://issues.apache.org/jira/browse/SPARK-25823?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16662660#comment-16662660
]
Dongjoon Hyun edited comment on SPARK-25823 at 10/24/18 6:38 PM:
-----------------------------------------------------------------
Ur, I think `collect` is the correct one as you can see as CTAS example. We
save with the last win entries.
BTW, [~cloud_fan]. I'm looking around this issue. I'll create another
improvement issue to fix `show` function for the following your comment.
{quote}BTW one improvement we can do is to remove duplicated map keys when
converting map values to string, to make it invisible to end-users.
{quote}
It's a regression introduced by SPARK-23023 at Spark 2.3.0. cc [~maropu].
{code:java}
scala> sql("SELECT map(1,2,1,3)").show
+---------------+
|map(1, 2, 1, 3)|
+---------------+
| Map(1 -> 3)|
+---------------+
scala> spark.version
res2: String = 2.2.2
{code}
was (Author: dongjoon):
BTW, [~cloud_fan]. I'm looking around this issue. I'll create another
improvement issue to fix `show` function for the following your comment.
bq. BTW one improvement we can do is to remove duplicated map keys when
converting map values to string, to make it invisible to end-users.
It's a regression introduced by SPARK-23023 at Spark 2.3.0. cc [~maropu].
{code}
scala> sql("SELECT map(1,2,1,3)").show
+---------------+
|map(1, 2, 1, 3)|
+---------------+
| Map(1 -> 3)|
+---------------+
scala> spark.version
res2: String = 2.2.2
{code}
> map_filter can generate incorrect data
> --------------------------------------
>
> Key: SPARK-25823
> URL: https://issues.apache.org/jira/browse/SPARK-25823
> Project: Spark
> Issue Type: Bug
> Components: SQL
> Affects Versions: 2.4.0
> Reporter: Dongjoon Hyun
> Priority: Blocker
> Labels: correctness
>
> This is not a regression because this occurs in new high-order functions like
> `map_filter` and `map_concat`. The root cause is Spark's `CreateMap` allows
> the duplication. If we want to allow this difference in new high-order
> functions, we had better add some warning about this different on these
> functions after RC4 voting pass at least. Otherwise, this will surprise
> Presto-based users.
> *Spark 2.4*
> {code:java}
> spark-sql> CREATE TABLE t AS SELECT m, map_filter(m, (k,v) -> v=2) c FROM
> (SELECT map_concat(map(1,2), map(1,3)) m);
> spark-sql> SELECT * FROM t;
> {1:3} {1:2}
> {code}
> *Presto 0.212*
> {code:java}
> presto> SELECT a, map_filter(a, (k,v) -> v = 2) FROM (SELECT
> map_concat(map(array[1],array[2]), map(array[1],array[3])) a);
> a | _col1
> -------+-------
> {1=3} | {}
> {code}
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]